Code Owns the Simulation, Jev Owns the Evaluation
Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
Abstract
Judgment models such as return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. succeeds when the right option can be judged from what the input describes, which we call evaluation. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on simulation (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, becomes an expert controller through its general evaluation ability.
Create a lesson
Related papers
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Sohyeon Kim, Yoonho Lee, Bo Liu et al.
VISTA: A Visual Harness for Reasoning in an Interactive World
Qiushi Han, Keya Hu, Linlu Qiu et al.
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España et al.
Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints
Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir
PyPottery: an AI-powered end-to-end suite for pottery processing and publication
Lorenzo Cardarelli
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
Arman Behnam, Binghui Wang