MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval
Rohan Pandey, Sunjae Kwon, Hong Yu
Abstract
Scientific discovery depends on finding prior literature that shapes what comes next. Existing retrieval systems optimize for relevance and popularity, often favoring central papers over less familiar works that later prove generative. We introduce MUSES, a million-instance benchmark for prospective intellectual-roots retrieval over a fixed 2.33M-paper corpus, with roughly 140K test instances per familiarity tier. To our knowledge, it is the first prospective benchmark at this scale with a shared retrieval task and author-confirmed paper-level root labels. Alongside it, CiteRoots pairs a scalable rhetorical layer over local citation text (LLM judge κ= 0.896 versus human gold) with a paper-level author-endorsed layer (n = 1,518 generative-inspiration pairs from 753 focal papers). MUSES organizes difficulty along two axes: a familiarity axis spanning CiteNext, CiteNew, and CiteNew-Isolated, and a functional axis spanning broad citations, rhetorical roots, and author-endorsed roots. Across 9 method classes, a lean multi-centroid retriever built on SPECTER2 is strongest. Hit@100 falls from 0.534 on CiteNext to 0.424 on CiteNew, 0.205 on rhetorical CiteNew, and 0.171 on author-endorsed CiteNew, a 3.1× decline. In a registered eight-lens full-test audit, roughly half of broad-tier test instances remain unsolved at K=1,000. Rhetorical role and author endorsement are distinct: the same judge agrees with endorsement at κ= 0.037. We release MUSES, both CiteRoots layers, and a distilled open companion judge for future work on prospective retrieval and intellectual roots.
Create a lesson
Related papers
Closed Forms and Synthetic Twins: Predicting Approximate Nearest Neighbor Recall from Embedding Statistics
Shmuel Herman
Two-Sided State-Space Models for Sequential Recommendation with Non-Random Multimodal Review Feedback
Ziwen Pan, Zihan Liang, Ruoxuan Xiong
MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval
Seokwon Song, Sohyeon Kim, Gunhee Kim
Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval
Shaowei Wei, Chong Huang, Songtao Fang et al.
Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster
Songtao Fang, Zihao Xu, Shaowei Wei et al.
Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval
Gaoming Zhang, Angqing Jiang, Jianchun Song et al.