Latent Mechanisms of Language Control in Multilingual Language Models
Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly
Abstract
Multilingual large language models can exhibit unintended code-switching -- unnecessarily alternating between languages during generation. We present a comparative study of three methods that identify language-controlling latents in cross-layer transcoders: activation value-based selection (ValSel), activation frequency-based selection (FreqSel), and LLM-generated latent annotation-based selection (AnnSel). To evaluate the efficacy of these methods in identifying language-controlling latents, we introduce two multilingual benchmarks that exhibit code-switching for fine-grained analysis of language steering across seven languages. Through targeted intervention experiments on Gemma-2-2B and Qwen3-4B, we find that all three methods effectively manipulate generation language, with FreqSel achieving the strongest overall performance, while AnnSel offering interpretable latent selection through explicit language annotations. A knock-out analysis suggests the methods select non-overlapping but each-functional latent subsets, indicating redundancy rather than a single canonical language direction. Code and data can be found at https://github.com/rm-3284/Latent-Mechanism-Multilingual.
Create a lesson
Related papers
Location-Aware Language Models via Secondary Embeddings
Gokul Srinivasagan, Munir Georges
(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement
Zach Studdiford, Kanishka Misra
Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation
Christos Nikolaos Zacharopoulos, Revekka Kyriakoglou, Chara Tsoukala et al.
From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education
Gokhan Dogru, Adrià Martín Mor
Two locked tests of phase-structure features for transition prediction
Abraham Chachamovits
Emotional Labor Strategy Preferences in LLM Personas
Mohammad Saim, Tianyu Jiang