Detecting Inconsistencies in Model Specifications with LLM-as-Verifier Reasoning
Zichen Xie, Mrigank Pawagi, Lize Shao, Yang Hu, Wenxi Wang
Abstract
Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet these specifications may themselves contain defects: two individually reasonable principles may prescribe incompatible behavior when applied to the same situation, leaving no response that satisfies both. Detecting such inconsistencies is challenging. Formalizing natural-language specifications risks losing subtle distinctions, while behavior-based testing cannot reliably distinguish specification defects from differences in model behavior. We introduce VeriSpec, the first approach to directly detect inconsistencies in model specifications by auditing the specification text itself. Our key insight is to preserve the specification in natural language while using an LLM as a verifier. VeriSpec extracts structured, context-aware rules, constructs a topic-guided graph to cluster behaviorally related rules at the same authority level, and applies LLM-as-verifier reasoning to detect inconsistencies. Applying VeriSpec to the OpenAI Model Spec, we extract 405 rules and manually validate five inconsistencies, all reported to its developers, who responded positively and have initiated internal discussions. Compared with five baselines, VeriSpec identifies the most validated inconsistencies, achieves the highest precision (38.5%), and incurs the lowest cost per validated inconsistency ($11.12). These results establish direct specification auditing as a practical complement to behavioral alignment evaluation, catching defects at the source before they shape any model. The code is available at https://github.com/HIPREL-Group/VeriSpec.
Create a lesson
Related papers
CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
Zheng Fang, Yongmin Li, Yichang Zhang et al.
Code Detectors Have a Half-Life: Obsolescence and Metric Illusions in LLM-Generated Code Detection
Alberick Euraste Djire
Architectural Degradation: How to Measure and to Remediate
Noman Ahmad, Ruoyu Su, Matteo Esposito et al.
Refactoring React Component Hierarchies to Eliminate Prop Drilling
Vangelis Gkinis, Vassilis E. Zafeiris
A Design Theory for AI-Assisted Software Development Derived from Christopher Alexander's Theory of Form
Chien-Tsun Chen, Yu Chin Cheng
CompProv Produces Machine Readable Graphs Encoding Microscopic Algebraic Provenance for Reproducible Computation
Minas Abramyan, Mohammed Alaa Ala'anzy, Nasir Saeed