Accepted · GroundLM @ EMNLP 2026
Siddhesh More, Kunal JadhavArizona State University
GroundLM 2026 Workshop @ EMNLP 2026 (LitTraceQA shared task), 2026
We present a stage-wise diagnostic of a literature-grounded question answering pipeline on LitTraceQA, a shared task that requires retrieving the right paper, locating the evidence within it, and generating a faithful answer. An oracle-substitution ladder, which hands the pipeline the gold paper and then the gold evidence, shows that evidence localization is the dominant remediable bottleneck: multiple-choice accuracy is 0.756 with the gold evidence value and 0.829 with matched page context, but 0.415 with predicted evidence. Cross-model hypothesis aggregation lifts paper-level F1 to 0.199, against 0.180 for single-model HyDE and 0.168 for BM25. We also report two infrastructure findings: the benchmark's page locators are keyed to a paper's canonical published version rather than its arXiv preprint, and many OpenReview PDFs sit behind a JavaScript bot challenge. Results use a 55-question dev set with bootstrap confidence intervals, and we frame the work as a resource and lessons-learned contribution rather than a leaderboard claim.