Do Multilingual Encoders Produce Language-Consistent Semantic IDs?
Abhinav Bohra, Anuj Bohra
Abstract
Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spanish, and Japanese, we test whether translations remain close to their English source, whether residual quantization is unusually sensitive to translation-induced movement, and whether multilingual or language-balanced quantizer fitting improves SID agreement. Multilingual E5 places translations measurably apart: under an English-heavy fit, a Japanese translation preserves the first SID code of its English counterpart in only 7.7% of cases, compared with 89.0% for an English rewording. Distance-matched product-directed controls produce nearly the same full-SID mismatch as translation, providing no evidence that the quantizer selectively amplifies language directions. Balancing the fitting mixture makes codebook use more uniform but further reduces cross-lingual prefix agreement: Spanish first-code consistency falls from 28.3% to 6.6%, while an English-only fit preserves it for 67.6% of Spanish translations. These results show that multilingual exposure and balanced codebook use alone do not guarantee language-consistent SIDs.
Create a lesson
Related papers
Optimizing Effective Training Time for Large-Scale Recommendation Systems
Mingming Ding, Ruilin Chen, Yuzhen Huang et al.
AgentWebRec: Compact Evidence Fusion over the Agent Web for Personalized Recommendation
Haoran Qiang, Guannan Liu, Liang Zhang et al.
From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search
Nikolai Zenovkin, Sebastian Björkqvist
Learning to structure data from user-generated thematic corpora
Elishay Avram, Oren Glickman, Elad Yom-Tov
RPTune: Learned Context Curation for LLM Catalog Search
Chuxuan Hu, Hejie Cui, Norman Huang et al.
CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG
Hyojeong Yun, Jueun Kim, Wook-Shin Han