CompLit: Scientific Literature Search Benchmarks Must Cover Implicit, Cumulative, and Unmet Needs
Published in arXiv 2026, 2026
CompLit is a scientific literature search benchmark spanning all eight arXiv domains. Researchers describe what they want in their own terms, which a paper may show only implicitly; their requirements grow as they read; and what they seek may lie where no paper yet exists. CompLit has one diagnostic setting for each of these needs. The strongest agents, Codex (GPT-5.6-Sol) and Claude Code (Opus 4.8), answer 86–87% of Standard queries, yet fall to 66–68% when requirements are implicit, lose another 6 and 15pp as they accumulate, and abstain on at most 53% of the queries no paper satisfies.
| Project page | Code | Dataset |
Recommended citation: Dayoon Ko, Jihyuk Kim, Soyeong Jeong, Young-Jun Lee, Dahyun Lee, Juyeon Kim, Gunhee Kim, Moontae Lee, Kyungjae Lee. (2026). "CompLit: Scientific Literature Search Benchmarks Must Cover Implicit, Cumulative, and Unmet Needs." arXiv preprint.
