Unmasking Face Embeddings: Reading, Rendering and Naming with Foundation Models
Fizza Rubab, Yiying Tong, Arun Ross
Abstract
Modern face recognition (FR) owes much of its success to deep neural networks that learn to extract compact identity embeddings from face images. These models are typically trained for identity discrimination, producing embeddings that are highly effective for biometric matching but largely opaque to semantic interpretation. In contrast, foundation models, pretrained on broad visual or vision--language tasks, provide rich interfaces for describing, retrieving, generating, and organizing visual content. This contrast raises a natural question: what capabilities become available when face embeddings from domain-specific FR models are made interoperable with foundation models? Building on recent work on embedding compatibility across models, we use simple pre-computed linear transformations, estimated from paired embeddings alone, to connect existing FR models with off-the-shelf foundation models. Once aligned with a foundation model, a face embedding can be 'unmasked' in multiple ways, without training or modifying either model: it can be read in natural language, enabling free-form text queries over a gallery of FR embeddings; rendered into a face image that recovers a person's appearance, using an unmodified diffusion decoder; and converted to a name, enabling identification even in the absence of an enrolled face gallery. In effect, one linear transformation turns an identity embedding into a rich embedding for web-scale foundation models. This interoperability exposes face embeddings as semantically and visually rich biometric representations, with direct implications for interpretability, retrieval, reconstruction, and template security.
Create a lesson
Related papers
Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT
Zhenyu Bu, Haoyan Ding, Chushu Shen et al.
SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling
Chad Wong, Sicheng Chen, Tianyi Zhang et al.
FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
Maya Moriya, Sigal Raab, Yael Vinker et al.
Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
Vida Adeli, Soroush Mehraban, Jacob Rommann et al.
Where Should Experience Live? Hierarchical Hebbian Memory for Continual Vision Transformers
Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz
TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening
Parham Hajishafiezahramini, Matthew Hamilton, Edward Kendall et al.