RAG that finds the right evidence—not just similar text.
A production retrieval system combining hybrid search, LLM reranking, grounded generation, and deliberate cost controls for enterprise knowledge workflows.
What made this
worth solving.
Enterprise documents are noisy, long, and filled with near-duplicate language. Basic vector search surfaced plausible passages, but not always the evidence a high-stakes answer actually needed.
Engineering choices,
not feature lists.
- 01
Combined semantic and keyword retrieval to preserve both meaning and exact business terminology.
- 02
Added LLM reranking and evidence-aware generation so the final answer stayed tied to the strongest sources.
- 03
Introduced prompt caching, token compression, and model routing to reduce cost without flattening answer quality.
What changed after
the system shipped.
The redesigned pipeline improved retrieval precision by approximately 35% while reducing LLM API cost by 40%. The same architecture patterns now inform production RAG and proposal-automation work.