How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation
Elitsa Yotkova, Violeta Kastreva, Petar Velkov, Hristo Boyanov, Dimitar Dimitrov, Ivan Koychev, Preslav Nakov
Abstract
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surface forms and may fail in qualitatively different ways, including incompleteness, contradiction, overgeneration, and endorsement of false premises. Existing judgment-based and similarity-based metrics often collapse these distinctions. We address this gap with three reusable contributions. First, we introduce a semantic correctness taxonomy that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content. Second, we release CAP-Correctness, an 8.8k-example benchmark spanning widely used QA datasets, and CAP-Statements, an 11k-example dataset for converting question-answer pairs into declarative statements for natural language inference (NLI) training and statement-based evaluation. Third, we introduce CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI. Under a monotonicity protocol testing whether metrics respect the taxonomy's intended ordering, CAP outperforms established baselines.
Create a lesson
Related papers
User Feedback Provides a Unique Signal that LLMs Can not Detect
Shachar Don-Yehiya, Leshem Choshen, Omri Abend
DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation
Vasileios Baltatzis, Mert Inan, Connor Gillis et al.
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
Yuling Shi, Zhensu Sun, Junsen Dong et al.
HyperStyler: Low-resource Authorship Style Transfer via Context-aware Style Navigation and Hypernetworks
Jongkyung Shin, Minguk Jeon, Chanwoo Park et al.
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Yuzhang Luo, Chenpeng Wang, Jianhui Chen et al.
Untangling the Mechanisms of Misleading Context in Medical Question Answering
Robin Linzmayer, Noémie Elhadad