Graph Evidence Is Not Enough: Diagnosing Native Decoder Use in Graph-Augmented LLMs
Xiaoyu Guo, Pengcheng Chen, Jiong Yu, Yi Lu, Yaohua Wang, Ziyang Li
Abstract
Graph-augmented large language models often assume that graph evidence produced by external computation and placed in the input can be used by the native decoder. We test this assumption with HopQA, a deliberately bounded diagnostic that asks for the shortest-hop distance between two query nodes. Because the answer is a small integer and the target is purely topological, failure cannot be dismissed as open-ended generation or ambiguous evaluation. Yet existing graph-augmented baselines still fail on this setting, showing that providing graph evidence is not the same as making it usable. We introduce an intervention triangle with three matched conditions: readable graph evidence, shuffled graph evidence, and no-graph input. This separates evidence inclusion, structural readability, and decoder-usable topology. Guided by this diagnosis, we present S2GE as an instance showing that diagnosis-driven interface design can improve native decoder usability. S2GE uses query-aware sampling, endpoint and proximity-based ordering, and structure-preserving alignment. Across DBLP, Biomedical, GoodReads, and PubMed, S2GE achieves strict exact-match scores of 36.5\%, 57.8\%, 76.6\%, and 52.0\%, improving over the strongest native-generation baseline by 53.5 points on average. The interventions further reveal harmful-shuffle, shuffle-robust, and no-graph-saturated regimes.
Create a lesson
Related papers
User Feedback Provides a Unique Signal that LLMs Can not Detect
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Vasileios Baltatzis, Mert Inan, Connor Gillis et al.
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Yuling Shi, Zhensu Sun, Junsen Dong et al.
HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks
Jongkyung Shin, Minguk Jeon, Chanwoo Park et al.
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Yuzhang Luo, Chenpeng Wang, Jianhui Chen et al.
Untangling the Mechanisms of Misleading Context in Medical Question Answering
Robin Linzmayer, Noémie Elhadad