Hybrid RAG Service

Role
Sole author
Type
Prototype, NDA
Focus
Retrieval & ranking

A full-stack retrieval app that runs full-text and vector search side by side in PostgreSQL, fuses the rankings and lets an LLM rerank the shortlist. Under NDA, so this covers the architecture only.

Architecture

Hybrid retrieval in PostgreSQL
Full-text ranking and pgvector cosine search run in parallel over the same table.
Reciprocal Rank Fusion
The two rankings are fused with RRF, with ties broken by vector rank.
LLM rerank
A deterministic, JSON-mode completion reorders the fused shortlist.
Document ingestion
Text is extracted from PDF and DOCX uploads, and an LLM turns it into structured query context.
Provider-agnostic LLM layer
One OpenAI-compatible client, switched between providers by configuration for chat and embeddings.
MCP client over HTTP
Raw JSON-RPC 2.0 requests with Server-Sent Events response parsing.
Sync job
An endpoint pulls records from an external API into the local store, with a mock fallback for development.

Main processing path

  1. document uploadPDF, DOCX
  2. context extractionLLM
  3. embeddingvectors
  4. full-text + vectorin parallel
  5. rank fusionRRF
  6. rerankLLM

Simplified, with generic component names. Highlighted steps use a model; the rest is software.

Stack

  • Python
  • FastAPI
  • PostgreSQL
  • pgvector
  • SQLAlchemy
  • React
  • TypeScript
  • Vite