ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Zikai Xiao, Zhenghong Lin, Qing Xu, Erik Cambria, Xiaoying Tang, Haoran Luo
Abstract
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval spans 23 disciplines and measures scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost through sequential streams and independent reset evaluation. Our framework repairs executable workflows through multi-turn interaction, converts re-execution-verified failure--success trajectories into linked Skill and Operator candidates, and retains an update only when source-task replay reproduces the repair and independent scientific tasks improve. Code is available at https://github.com/beita6969/ScienceClaw.
Create a lesson
Related papers
Sherpa: Teaching LLMs to Teach Adaptively
Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang et al.
Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
Zewei Zhou, Rachel Luo, Yulong Cao et al.
Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Egor Pakhomov, Erik Nijkamp
WorldSolver: Can LLM Agents Simulate the Physical Dynamics via Solver Generation?
Siru Jiang, Yongzhe Lyu, Shuo Lu et al.
nanoMuse: An Open-Source Personal Agent for Every Device You Own
Guangyi Liu, Yong Liu, Jiangning Zhang