Inside the provider-agnostic, three-tier fallback architecture that keeps enterprise AI workflows running when models falter. Every client talks to one service, every heavy task runs on one queue, and every AI-backed capability inherits the same reliability ladder.
~60%
less AI code surface after unifying five provider integrations behind one service
5
production AI workflows on one shared reliability pattern
0
new infrastructure for v1 of the RAG tier, built on BullMQ, Redis and pgvector
“The AI makes the system smart; the fallback ladder makes it something you can depend on during business hours, every day.”
01 · The problem
LLMs fail in production
Third-party LLM APIs hit unpredictable rate limits (HTTP 429), regional outages (HTTP 503), latency spikes of ten seconds or more, and strict token budgets. A platform that couples core operations directly to raw LLM calls stops the moment one upstream provider does, taking hiring pipelines and recruiter consoles with it.
02 · The approach
One service, one ladder
A single Unified AI Service, backed by idempotent BullMQ and Redis queues, runs every request down a strict three-tier ladder: automatic multi-provider retry and failover, then tenant-scoped RAG retrieval, then a deterministic rule-based floor. Users never see a 500 error.
03
Architecture
Recruiter console
ATS / HRMS
Careers portal
Partner APIs
↓ single provider contract · NestJS guard and ingestion controller
Unified AI Service
One entry point for every AI request, one provider contract and telemetry hub
↓ async offload
Async work queue · BullMQ + Redis
Idempotent, retry-safe workers with exponential backoff
↓ five shared capabilities
Smart parsing
documents and OCR
Candidate scoring
explainable fit
Interview questions
role-aware
Hiring verdict
evidence-backed
Assistant
grounded chat
↓ every capability inherits the ladder
Reliability ladder
AI answer → confidence and latency check → RAG grounded fallback → rule-based floor
Grounded knowledge (RAG)
Vector retrieval over the tenant's own jobs and candidates
Model providers
Gemini · OpenAI · Anthropic, with runtime failover
Figure 1 · Every client talks to one service. Every heavy task runs on one queue.
Write path · indexing, async on the BullMQ worker pool
{{ s.n }}
{{ s.t }}
{{ s.d }}
Read path · retrieval and grounded generation
{{ s.n }}
{{ s.t }}
{{ s.d }}
If the model provider is down: the read path returns ranked snippets instead. No generation is needed, and answers stay grounded and cited.
Right to erasure. Deleting a candidate or tenant triggers an atomic cascade that purges the relational record, the MeiliSearch entry, raw S3/GCS payloads and every derived 768-dimension embedding.
Figure 2 · The read path degrades to grounded snippets, never an error.