Hybrid RAG Service
A full-stack retrieval app that runs full-text and vector search side by side in PostgreSQL, fuses the rankings and lets an LLM rerank the shortlist. Under NDA, so this covers the architecture only.
Architecture
- Hybrid retrieval in PostgreSQL
- Full-text ranking and pgvector cosine search run in parallel over the same table.
- Reciprocal Rank Fusion
- The two rankings are fused with RRF, with ties broken by vector rank.
- LLM rerank
- A deterministic, JSON-mode completion reorders the fused shortlist.
- Document ingestion
- Text is extracted from PDF and DOCX uploads, and an LLM turns it into structured query context.
- Provider-agnostic LLM layer
- One OpenAI-compatible client, switched between providers by configuration for chat and embeddings.
- MCP client over HTTP
- Raw JSON-RPC 2.0 requests with Server-Sent Events response parsing.
- Sync job
- An endpoint pulls records from an external API into the local store, with a mock fallback for development.
Main processing path
- document uploadPDF, DOCX
- context extractionLLM
- embeddingvectors
- full-text + vectorin parallel
- rank fusionRRF
- rerankLLM
Simplified, with generic component names. Highlighted steps use a model; the rest is software.