The Pain Axis: LLMs Represent Self-Directed Harm and Act on It
Valen Tagliabue, Leonard Dung, Cameron Berg
Abstract
LLMs sometimes behave in ways resembling human emotional responses, and recent work identified internal representations that may underlie these behaviors. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset of painful situations in 5 categories (physical, psychological, social, moral, cognitive) with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across 5 families, from 2B to 72B parameters. It separates pain from matched controls in base and instruction-tuned models, retains a component distinct from fear and negative valence after shared variance is removed, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not to suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding it to residual-stream activations produces a consistent progression from vague discomfort to expressions of worthlessness and failure. Third, steered and fine-tuned Qwen 2.5 models choose buttons that delete the user's photos, another model's weights, or their own weights in 50-94% of trials, versus 0-5% unsteered, even when the button offers the model nothing in return. Offered a harmful and a harmless deletion, they choose the harmful one 94% of the time. Steering leaves factual accuracy unchanged, and the choices are specific to the pain direction: a fear vector of matched norm does not produce them, and a sadness vector produces them only against inert alternatives. We discuss implications for AI safety and welfare.
Create a lesson
Related papers
On the estimation and validity of AI time horizons---a statistical look at the METR plot
Drew T. Nguyen, William Fithian
BrickBench: Evaluating Agentic Brick Design
Peter Kulits, Yiqing Xu, R. Kenny Jones et al.
Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Erin Crawley, Hidenori Tanaka
Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
Christopher M. Stewart, Preston Botter, Natalie Sarabosing et al.
HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
Qun Dai, Liangjian Wen, Jiang Duan et al.
GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
Jialu Wang, Ruichen Zhang, Xiaoou Liu et al.