Skip to content

Simulate, record, verify: A language-portable framework for muscle-grounded articulatory QA (extended version)

Seungho Eum, Shantong Sun, Unsang Park

cs.CVarXiv:2608.23137

Abstract

Articulatory data describe what the tongue looks like, but geometry alone does not explain the muscle-level why behind a configuration or how to move it toward a target posture. We present a simulator-based framework that constructs this supervision from controlled biomechanical inputs. Each simulated configuration is stored with its generating muscle state and geometric properties in a structured fact record. Gold answers are derived before linguistic realization, and every naturalized output is checked against its source record, so the same records support new languages and question types without re-simulation. We instantiate the framework as 3DTongueQA using the ArtiSynth Badin tongue model. From 295,115 meshes, we construct 891,156 record-checked QA instances per language, and the checker detects over 96.9% of injected corruptions. Korean and Spanish realizations and a new question type are added from the same records, confirming portability. A SpiralNet++-Qwen3 probe reaches 62.9 Muscle EM, above 44.0 for nearest-neighbor transfer and 7.2 for a zero-shot commercial LLM given the same mesh, while mesh shuffling drops it to near zero. The supervision is learnable, useful beyond retrieval, and grounded in paired geometry. It reflects simulator-defined states rather than measured physiology. Code, templates, and QA sets are available at https://github.com/esh0504/muscle-grounded-qa.

Create a lesson