Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
Rana Muhammad Usman, Dominic Williamson
Abstract
Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.
Create a lesson
Related papers
Towards stratified sampling for redistricting plans
Zijian Wang, Gregory J. Herschlag, Joon-Hyeok Yim et al.
BanglaShop-CRS: A User-Centric Bangla Dataset for Conversational Recommendation
Tabia Tanzin Prama, Christopher M. Danforth, Peter Sheridan Dodds
Ensuring proportionality: a logical model for compensatory seats added to multi-member constituencies
Frederik Ravn Klausen
Modelling opinion dynamics during crises as complex contagion with feedback
Junxiang Huang, Mikhail Prokopenko
Topological Uncertainty and Higher-Order Interactions in Spatial Networks
Domenico Pomarico, Alessandro Fania, Gabriel Ramirez Sanchez et al.
Shannon entropy and complex network community detection to study electoral coalition behaviour at municipal scale
Andrea Lo Sasso