JUH Jawad Ul Hadi
Résumé
← Back to selected work

Case study · Resilience architecture

Designing for AI failure

RAG and BullMQ fallback architecture

Inside the provider-agnostic, three-tier fallback architecture that keeps enterprise AI workflows running when models falter. Every client talks to one service, every heavy task runs on one queue, and every AI-backed capability inherits the same reliability ladder.

~60%

less AI code surface after unifying five provider integrations behind one service

5

production AI workflows on one shared reliability pattern

0

new infrastructure for v1 of the RAG tier, built on BullMQ, Redis and pgvector

3

fallback tiers: queue retry, grounded RAG answer, rule-based floor

“The AI makes the system smart; the fallback ladder makes it something you can depend on during business hours, every day.”

01 · The problem

LLMs fail in production

Third-party LLM APIs hit unpredictable rate limits (HTTP 429), regional outages (HTTP 503), latency spikes of ten seconds or more, and strict token budgets. A platform that couples core operations directly to raw LLM calls stops the moment one upstream provider does, taking hiring pipelines and recruiter consoles with it.

02 · The approach

One service, one ladder

A single Unified AI Service, backed by idempotent BullMQ and Redis queues, runs every request down a strict three-tier ladder: automatic multi-provider retry and failover, then tenant-scoped RAG retrieval, then a deterministic rule-based floor. Users never see a 500 error.

03

Architecture

Recruiter console
ATS / HRMS
Careers portal
Partner APIs

↓ single provider contract · NestJS guard and ingestion controller

Unified AI Service

One entry point for every AI request, one provider contract and telemetry hub

↓ async offload

Async work queue · BullMQ + Redis

Idempotent, retry-safe workers with exponential backoff

↓ five shared capabilities

Smart parsing

documents and OCR

Candidate scoring

explainable fit

Interview questions

role-aware

Hiring verdict

evidence-backed

Assistant

grounded chat

↓ every capability inherits the ladder

Reliability ladder

AI answer → confidence and latency check → RAG grounded fallback → rule-based floor

Grounded knowledge (RAG)

Vector retrieval over the tenant's own jobs and candidates

Model providers

Gemini · OpenAI · Anthropic, with runtime failover

Multi-tenant security: vaulted API secrets · tenant-scoped vector queries · audit logging · GDPR erasure cascade

Figure 1 · Every client talks to one service. Every heavy task runs on one queue.

Write path · indexing, async on the BullMQ worker pool

  1. {{ s.n }}

    {{ s.t }}

    {{ s.d }}

Read path · retrieval and grounded generation

  1. {{ s.n }}

    {{ s.t }}

    {{ s.d }}

If the model provider is down: the read path returns ranked snippets instead. No generation is needed, and answers stay grounded and cited.

Right to erasure. Deleting a candidate or tenant triggers an atomic cascade that purges the relational record, the MeiliSearch entry, raw S3/GCS payloads and every derived 768-dimension embedding.

Figure 2 · The read path degrades to grounded snippets, never an error.
  1. {{ s.t }}

    {{ s.d }}

{{ t.tier }}

{{ t.when }}

{{ t.path }}

Figure 3 · Retry stays inside the queue, RAG catches provider failure, rules catch empty retrieval.
{{ l.name }} {{ l.state }}

{{ l.desc }}

Telemetry {{ latency }}
{{ g.t }} {{ g.level }} {{ g.msg }}

Simulation · pick a failure scenario to watch the ladder step down.

04

Design guarantees

{{ g.t }}

{{ g.d }}

05

Implementation

UnifiedAIService.ts · TypeScript / NestJS

export class UnifiedAIService {
  async executeWithReliabilityLadder<T>(
    req: AIRequestPayload,
    tenantContext: TenantSecurityContext
  ): Promise<AIResponse<T>> {
    // Tier 1: multi-provider LLM with exponential retry
    try {
      return await this.dispatchPrimaryLLM(req, tenantContext);
    } catch (llmErr) {
      this.logger.warn(`Tier-1 LLM failed: ${llmErr.message}. Escalating to Tier-2 RAG.`);
    }

    // Tier 2: tenant-scoped grounded vector retrieval
    try {
      const vectorMatches = await this.vectorStore.searchByTenant({
        tenantId: tenantContext.tenantId,
        query: req.queryText,
        threshold: 0.78
      });
      if (vectorMatches.length > 0) {
        return this.formatGroundedSnippets(vectorMatches, 'TIER_2_RAG');
      }
    } catch (ragErr) {
      this.logger.warn(`Tier-2 RAG empty/failed: ${ragErr.message}. Escalating to Tier-3 floor.`);
    }

    // Tier 3: deterministic rule-based heuristics and OCR floor
    return this.ruleEngine.extractHeuristicData(req.rawDocument, 'TIER_3_RULE_FLOOR');
  }
}

Want the pattern on your platform?

Contact me on Gravatar ↗ See services
© 2026 Jawad Ul Hadi · Backend Lead / Architect “The best code is never rewritten, because it is flexible enough to evolve with the business.”