Location-Aware Language Models via Secondary Embeddings
Gokul Srinivasagan, Munir Georges
Abstract
Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational semantics, leading to suboptimal representations of place names and spatial entities. In this work, we propose a lightweight, model-agnostic approach for injecting geo-spatial awareness into pretrained embeddings without modifying the tokenizer or requiring costly retraining. Our method augments input representations with structured geographic signals by combining location names with their corresponding latitude and longitude, and employs a location-focused masking to better align textual representations with real-world spatial relationships. This design allows the model to incorporate geo-spatial context while preserving existing semantic and syntactic knowledge. Experimental results demonstrate substantial improvements in geo-spatial alignment while maintaining comparable performance on standard NLP benchmarks such as GLUE. The method is computationally efficient, requiring only minutes of additional training, and generalizes across multiple model architectures and scales.
Create a lesson
Related papers
(V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement
Zach Studdiford, Kanishka Misra
Late Transformer Layers Recode Syntax Canonically: Evidence from Greek Scrambling and Cross-Layer Generalisation
Christos Nikolaos Zacharopoulos, Revekka Kyriakoglou, Chara Tsoukala et al.
From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education
Gokhan Dogru, Adrià Martín Mor
Two locked tests of phase-structure features for transition prediction
Abraham Chachamovits
Latent Mechanisms of Language Control in Multilingual Language Models
Ryo Mitsuhashi, Sabri Boughorbel, Majd Hawasly
Emotional Labor Strategy Preferences in LLM Personas
Mohammad Saim, Tianyu Jiang