Back

Accepted · GroundLM @ EMNLP 2026

Retrieve, Locate, Generate: An Oracle-Substitution Diagnostic for Literature-Grounded QA

Siddhesh More, Kunal JadhavArizona State University

GroundLM 2026 Workshop @ EMNLP 2026 (LitTraceQA shared task), 2026


Summary

We present a stage-wise diagnostic of a literature-grounded question answering pipeline on LitTraceQA, a shared task that requires retrieving the right paper, locating the evidence within it, and generating a faithful answer. An oracle-substitution ladder, which hands the pipeline the gold paper and then the gold evidence, shows that evidence localization is the dominant remediable bottleneck: multiple-choice accuracy is 0.756 with the gold evidence value and 0.829 with matched page context, but 0.415 with predicted evidence. Cross-model hypothesis aggregation lifts paper-level F1 to 0.199, against 0.180 for single-model HyDE and 0.168 for BM25. We also report two infrastructure findings: the benchmark's page locators are keyed to a paper's canonical published version rather than its arXiv preprint, and many OpenReview PDFs sit behind a JavaScript bot challenge. Results use a 55-question dev set with bootstrap confidence intervals, and we frame the work as a resource and lessons-learned contribution rather than a leaderboard claim.

literature-grounded QAevidence localizationoracle substitutionretrieval-augmented generation