Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation
Jinyoung Kim, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, Lu Wang
Abstract
Natural-language critiques provide supervision beyond scalar rewards for non-verifiable generation, which lacks deterministic verifiers. In critique-guided refinement, a critic gives feedback on an initial response and an actor revises it. However, final revision quality does not reveal whether the critique was actually useful: a capable actor may improve without following the feedback, while valid feedback may fail if the actor cannot execute it. We frame critique as actor-conditioned revision guidance, where usefulness depends on whether the feedback helps the target actor address the intended weakness. We introduce TAIScore (Targeted Actionable Improvement Score), a reward that evaluates the instruction, initial response, critique, and revision together, assessing whether the critique targets a real weakness, whether the actor follows it, and whether the intended aspect improves. We use this reward to train an actor-tailored critic with GRPO, and use critique-guided refinements to construct DPO preference pairs for the actor, forming a co-evolving critic-actor loop where the critic adapts to the actor's changing capability. Experiments show that an 8B critic trained with TAIScore outperforms both a zero-shot 120B critic and critics trained with outcome-only or critique-only reward signals. Co-evolving the critic and actor further improves performance, suggesting that effective critique supervision should adapt as the actor changes.
Create a lesson
Related papers
When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models
Jiaqi Wei, Xiang Zhang, Yuejin Yang et al.
Quantitative Evidence Mining for Plausibility-Aware Biomedical AI
Negin Sadat Babaiha, Stefan Geissler, Marie-Christine Simon et al.
Using Grounded Theory for Agent Behavior Analysis at Scale
Zhuoran Lu, Yangyang Yu, Zhuoyan Li et al.
Kathleen Remembers: Length-Invariant One-Shot Recall Without Attention
George Fountzoulas
Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation
Minsoo Song, Chanwoo Kim, Sugyeong Eo et al.
Auditing MCQA Benchmarks through Probability Landscapes
Minsoo Song, Chanjun Park