Project 12
A question-answering system over the last three annual reports (SEC Form 10-K) of 25 large US companies: 74 filings split into 25,496 chunks. It answers with citations, or says the answer is not in the documents. The point was to measure which parts of a RAG pipeline earn their cost, instead of assuming they do.
The pipeline ingests filings, runs hybrid search with company and year filters, reranks candidates with a cross-encoder, and routes questions about reported figures to a structured fact store instead of text retrieval. Passages and facts live in Postgres with pgvector, behind a FastAPI backend and a React frontend, with Terraform for AWS (validated, not applied).
How does a question flow through the system?
Questions about reported figures are routed to the structured fact store; text questions go through filtered hybrid search and reranking. Every answer carries citations, or abstains.
Evaluation used 180 numeric questions, 152 verified text questions (39 multi-hop) and 39 abstain questions, plus paraphrased, near-miss and injection tiers, with a free 30B model answering. Reranking was the stage that mattered: a paired bootstrap puts its gain at +2.7 to +15 points, while the rest of the full system is not distinguishable from reranking alone.
The full system answers at p50 2.26 s (p95 2.87 s) against 1.13 s for naive RAG, with one request at a time and fresh caches. Putting the whole context in the prompt instead took about 86k tokens and 95 s per question. pgvector HNSW at 1,024-d matched the in-memory 4,096-d index (recall@8 0.742 against 0.738 on 244 questions). A second domain, 28 FDA drug labels with 69 verified questions, reached recall@8 of 0.99-1.0.
Numeric questions answered from the fact store scored 100%, but the store and the ground truth both come from the same XBRL data, so that number is circular. The abstention score of 98% on near-miss questions comes from a rule shaped by that tier's own errors, so it is a development-set figure. The injection test used five templates and both configurations held (40 of 40 against 39 of 40), which does not distinguish them. The FDA result is saturated and measured retrieval only. Judging used LLMs (201 of 209 agreement between judges) with no human check, only 10 questions per tier backed the long-context comparison, and no paid model was run.
What I Learned
Tech Stack