Where LLMs Fail with Visualization DSLs
Chang Han, Andrew McNutt, Katherine Isaacs
Abstract
As LLMs take up the role of authoring charts using visualization domain-specific languages (DSLs), the human constraints that shaped those languages may no longer apply, as what is easy for a person is not necessarily easy for a model. To understand how LLMs might work better with DSLs, we explore where and how they fail with current DSL designs. We evaluate 10 JSON-style visualization DSLs with 41 tasks across 3 LLMs, then assess the generated specifications with JSON and rendering checks, and qualitative coding of failed cases. Analyzing how this specification generation process fails, we identify four recurring failure patterns, link each to specific DSL features, and discuss design considerations for future DSL designs.
Create a lesson
Related papers
XAI Evaluation Cards: A Practical Method for Designing Human-Centred XAI Evaluations
Kristýna Sirka Kacafírková, Ivania Donoso-Guzmán, Denis Parra et al.
Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI
Xiaotian Su, Laura Rimell, Jiazheng Li et al.
Scaling Peer Assessments: An Integrity Report from a Large Engineering Internship
Jinal Gupta, Pavani Ayinampudi, Aditya B. M. V. et al.
LeanSide: A Formally Verified Co-Reasoning System for Natural-language Proofs
Chenjun Guo, Manooshree Patel, Arnav Mehta et al.
From Images to Tasks: Characterizing Multimodal LLM Interactions in the Wild
Jinyi Ye, Scott Counts, Gaurav Verma et al.
Sensing Instability, Adapting the Scene: A Real-Time Movement-Smoothing Design Framework for Stable VR Locomotion
Ramisa Fariha Joyee, M. Rasel Mahmud