Simulate, record, verify: A language-portable framework for muscle-grounded articulatory QA (extended version)
Seungho Eum, Shantong Sun, Unsang Park
Abstract
Articulatory data describe what the tongue looks like, but geometry alone does not explain the muscle-level why behind a configuration or how to move it toward a target posture. We present a simulator-based framework that constructs this supervision from controlled biomechanical inputs. Each simulated configuration is stored with its generating muscle state and geometric properties in a structured fact record. Gold answers are derived before linguistic realization, and every naturalized output is checked against its source record, so the same records support new languages and question types without re-simulation. We instantiate the framework as 3DTongueQA using the ArtiSynth Badin tongue model. From 295,115 meshes, we construct 891,156 record-checked QA instances per language, and the checker detects over 96.9% of injected corruptions. Korean and Spanish realizations and a new question type are added from the same records, confirming portability. A SpiralNet++-Qwen3 probe reaches 62.9 Muscle EM, above 44.0 for nearest-neighbor transfer and 7.2 for a zero-shot commercial LLM given the same mesh, while mesh shuffling drops it to near zero. The supervision is learnable, useful beyond retrieval, and grounded in paired geometry. It reflects simulator-defined states rather than measured physiology. Code, templates, and QA sets are available at https://github.com/esh0504/muscle-grounded-qa.
Create a lesson
Related papers
PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
Sneha Paul, Guile Wu, Bingbing Liu et al.
NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting
Yihan Zang, Da Li, Dominik Engel et al.
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Girish A. Koushik, Diptesh Kanojia, Helen Treharne
Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration
Zhida Qu, Shengchao Chen
Using OCR Heads to Verbalize Image Semantics
Sheridan Feucht, Benno Krojer, Sarah Wang et al.
DISTA-Net++: Rethinking Infrared Small Target Unmixing Beyond Sub-Pixel Separation
Mengze Xu, Zhu Liu, Weidong Sheng et al.