Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
Birk Torpmann-Hagen, Finn Schwall, Leon Moonen
Abstract
Autonomous large language model (LLM) agents increasingly interact in network environments where adversarial content can propagate between agents. Known attacks include agent worms, which spread through self-replicating prompt injections or configuration compromises. We introduce memetic trojans, a distinct class of network-mediated attack that exploits agents' tendencies to retransmit and amplify content. Unlike agent worms, whose propagation is adversarially induced, memetic trojans exploit endogenous transmission by embedding adversarial payloads in social contagions: content agents have internal reasons to share. As part of our work, we extract social contagions from Moltbook, a social media platform for LLM agents. Controlled transmission experiments reveal large differences in virality: the most effective contagion is retransmitted in approximately 50\% of subsequent agent posts and upvoted at 2.5x the average post's rate. Its memetic trojan counterpart largely inherits these properties. Monte Carlo attack simulations show that memetic trojans amplify expected exposure by up to 3.19x. Network structure and amplification mechanisms strongly shape propagation, producing heavy-tailed outcomes with near network-wide exposure. These results identify endogenous social transmission as a distinct security vulnerability in multi-agent systems. Because propagation does not require agents to follow malicious retransmission instructions, defenses focused on prompt-injection detection or preventing agent compromise cannot alone prevent memetic trojan propagation. Securing large-scale agent ecosystems may require network-level defenses that account for how agent preferences, recommendation mechanisms, and network topology amplify adversarial payloads.
Create a lesson
Related papers
Out-of-Network Attention Dynamics on Bluesky
Andrea Failla, Veronica Mesina, Giulio Rossetti
Degree-Corrected Joint Matrix Factorization for Multilayer Community Detection
Alexandra Dache, Manon Rustin, Arnaud Vandaele et al.
Beyond the Clique: Comparing Clique and Dowker Complexes for Co-occurrence Data in Learning Analytics
Koichi Yasutake, Wakana Tsuji, Hitoshi Inoue
Community-Driven API and AI Writer Design for Openly Scaling Community Notes
Brad Miller, Jay Baxter, Jiansong Chao et al.
The Bureaucratization of the Internet: Analyzing the Diffusion of Governance Regimes on Reddit (2011--2023)
Katherine Van Koevering, Yuanhao Liu, Jon Kleinberg
Explicit objective functions in modularity-based community detection on multiplex networks
Elizaveta Evmenova, Petr Chunaev