WorkWorlds: An Infrastructure for Evaluating AI Agents on Workplace Tasks
Yining Hua, Levi Lian
Abstract
Many knowledge-work benchmarks are constructed around individual tasks, with the context needed for each task selected together with or after the task has been specified. This design measures performance on workplace-like tasks in an environment assembled for the task. When task specification guides which context is selected, the evaluation can encode task information into the environment and pre-complete part of the information-localization work that workplace performance normally requires. We introduce WorkWorlds, an evaluation infrastructure that separates organizational state from task specification. A world first fixes a revision, date, and employee seat and materializes the organizational state that employee can access; tasks are introduced only afterward. We implement WorkWorlds in a primary synthetic pharmaceutical company with 8 measured tasks across 6 employee seats, and construct additional organizational worlds. Across 192 matched evaluations, task-level curation increased evidence access by 17.6 percentage points, from 72.8% to 90.4%, and criterion pass by 8.7 points, from 68.0% to 76.7%, while pass conditional on evidence access remained nearly unchanged; most of the measured difference occurred before the agent reached sufficient evidence.
Create a lesson
Related papers
Sherpa: Teaching LLMs to Teach Adaptively
Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang et al.
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Zewei Zhou, Rachel Luo, Yulong Cao et al.
Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Egor Pakhomov, Erik Nijkamp
WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
Siru Jiang, Yongzhe Lyu, Shuo Lu et al.
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Guangyi Liu, Yong Liu, Jiangning Zhang