Citation-grounded RAG with a measurable eval loop. Hybrid search, optional rerank, per-citation verification, and a golden question set that runs against the live API.
The brief.
Citation-grounded RAG with a measurable eval loop. Hybrid search, optional rerank, per-citation verification, and a golden question set that runs against the live API.
- What works today
- LlamaIndex hybrid search with optional Cohere rerank, served from a FastAPI service.
- Every answer returns citations, and each citation is checked against its source passage.
- A golden question set and a live eval script that writes a timestamped report against the hosted API.
- What I owned
- retrieval pipeline, citation verification, eval harness, API, deploy.
- Known limits
- One in-memory index per process; a restart loses it unless auto-seed or an external store is used.
- Citation verification runs as sequential LLM calls, so it adds latency.
- The golden set covers a small demo corpus, not a client's documents.




