Back
FilingLens: Cited Q&A over SEC Filings

Project 12

FilingLens: Cited Q&A over SEC Filings


A question-answering system over the last three annual reports (SEC Form 10-K) of 25 large US companies: 74 filings split into 25,496 chunks. It answers with citations, or says the answer is not in the documents. The point was to measure which parts of a RAG pipeline earn their cost, instead of assuming they do.

The Approach

The pipeline ingests filings, runs hybrid search with company and year filters, reranks candidates with a cross-encoder, and routes questions about reported figures to a structured fact store instead of text retrieval. Passages and facts live in Postgres with pgvector, behind a FastAPI backend and a React frontend, with Terraform for AWS (validated, not applied).

How does a question flow through the system?

  1. Question
    plus company and year filters
  2. Route
    figures to the fact store, text to search
  3. Hybrid search
    BM25 + dense over pgvector
  4. Cross-encoder rerank
    reorders the short list
  5. Answer
    cited, or 'not in these documents'

Questions about reported figures are routed to the structured fact store; text questions go through filtered hybrid search and reranking. Every answer carries citations, or abstains.

The Results

Evaluation used 180 numeric questions, 152 verified text questions (39 multi-hop) and 39 abstain questions, plus paraphrased, near-miss and injection tiers, with a free 30B model answering. Reranking was the stage that mattered: a paired bootstrap puts its gain at +2.7 to +15 points, while the rest of the full system is not distinguishable from reranking alone.

The full system answers at p50 2.26 s (p95 2.87 s) against 1.13 s for naive RAG, with one request at a time and fresh caches. Putting the whole context in the prompt instead took about 86k tokens and 95 s per question. pgvector HNSW at 1,024-d matched the in-memory 4,096-d index (recall@8 0.742 against 0.738 on 244 questions). A second domain, 28 FDA drug labels with 69 verified questions, reached recall@8 of 0.99-1.0.

The Honest Parts

Numeric questions answered from the fact store scored 100%, but the store and the ground truth both come from the same XBRL data, so that number is circular. The abstention score of 98% on near-miss questions comes from a rule shaped by that tier's own errors, so it is a development-set figure. The injection test used five templates and both configurations held (40 of 40 against 39 of 40), which does not distinguish them. The FDA result is saturated and measured retrieval only. Judging used LLMs (201 of 209 agreement between judges) with no human check, only 10 questions per tier backed the long-context comparison, and no paid model was run.

What I Learned

  • Measure each stage on its ownAdding stages one at a time with a paired bootstrap showed reranking carried the accuracy gain and the rest of the pipeline could not be told apart from it.
  • Check whether your ground truth is independentA 100% numeric score looked like a win until it was clear the answers and the grader read the same XBRL source. Saying so on the page matters more than the number.
  • Say which numbers were tuned onThe 98% near-miss abstention came from a rule adjusted against that very tier, so it is reported as a dev-set figure rather than a held-out result.

Tech Stack

PythonFastAPIReactTypeScriptPostgreSQLpgvectorDockerTerraform