Measurement Validity in LLM Cultural Alignment
An Duy Nguyen, Muhammad Aurangzeb Ahmad
Abstract
Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.
Create a lesson
Related papers
Does the Power-Law Advantage in the Tails Outweigh the Global q-Gaussian Description?
Eduardo Boor, Roberto da Silva, Joao Carlos Schmitt de Siqueira et al.
The Mechanics of Democratic Dominance: A System Dynamics Paradigm for Dynamic Consent Engineering
Muhammad Sukri Bin Ramli
Peer Review at Capacity: An editor's view
Yamir Moreno
Open, Reproducible Per-Bidding-Zone Carbon Intensity for Nordic Power Systems
Eirik Botten Nicolaysen
ButterMamba: Butterworth-Enhanced Spatial-Temporal Mamba for Efficient Traffic Flow Prediction
Limiao Zhang, Yuhui Lu, Jie Gao et al.
Integrating adaptive human behavior into epidemic models with large language models
Yicheng Mao, Haoyang Li, Rob Deardon et al.