Large language models simulate intersectional synthetic identities with a budget of one to two dimensions
Virgile Rennard, Christos Xypolopoulos
Abstract
Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.
Create a lesson
Related papers
Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
Abbas M. Rabiu, Abdulrazaq A. Zubair, Um-mulkhairi Ibrahim et al.
Could Underwater Data Centers Pose a Risk to AI Treaty Verification?
James Teague, Ashmita Rajmohan, Yannick Muehlhaeuser
Control-Theoretic Content Moderation
Benedetta Tessa, Serena Tardelli, Marco Avvenuti et al.
"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations
Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante et al.
Understanding AI Provider Recommendations in Local Service Markets
Hazem Ibrahim, Yasir Zaki
Toward a Time-Aware Assessment Framework for the Carbon Cost of AI-Enabled Decarbonization
Chenrui Xu, Burcu Akinci, Christopher McComb