From Images to Tasks: Characterizing Multimodal LLM Interactions in the Wild
Jinyi Ye, Scott Counts, Gaurav Verma, Kate Lytvynets, Weiwei Yang
Abstract
Multimodal large language models (LLMs) increasingly integrate vision and text, yet how people use them in natural settings remains underexplored. We seek to answer the question: when users upload images, what tasks are they trying to accomplish? Analyzing over 40,000 de-identified image-upload conversations from Microsoft Copilot, we characterize real-world multimodal use through a hierarchical framework of ten capabilities, spanning perception, cognition, and generation. First, we characterize the distribution and composition of these capabilities, finding that the majority of image-upload tasks involve multiple capabilities. Second, we find that multimodal use spans a broader and more diverse task space than text-only interactions, with asymmetric coverage and task classes that rely on cross-modal grounding. Third, mapping observed capability demand onto 253 existing benchmarks reveals uneven alignment between benchmark coverage and real-world use: benchmarks concentrate on perception and reasoning toward fixed answers, while common workflows involving text, code, and data generation remain comparatively undertested. We validate our taxonomy and findings on an independent ChatGPT dataset. Our results provide a large-scale empirical characterization of what users seek to accomplish with multimodal LLMs and highlight opportunities for benchmark design grounded in observed user demand.
Create a lesson
Related papers
XAI Evaluation Cards: A Practical Method for Designing Human-Centred XAI Evaluations
Kristýna Sirka Kacafírková, Ivania Donoso-Guzmán, Denis Parra et al.
Where LLMs Fail with Visualization DSLs
Chang Han, Andrew McNutt, Katherine Isaacs
Who Thinks First? Designing Productive Friction with Engage-to-Unlock GenAI
Xiaotian Su, Laura Rimell, Jiazheng Li et al.
Scaling Peer Assessments: An Integrity Report from a Large Engineering Internship
Jinal Gupta, Pavani Ayinampudi, Aditya B. M. V. et al.
LeanSide: A Formally Verified Co-Reasoning System for Natural-language Proofs
Chenjun Guo, Manooshree Patel, Arnav Mehta et al.
Sensing Instability, Adapting the Scene: A Real-Time Movement-Smoothing Design Framework for Stable VR Locomotion
Ramisa Fariha Joyee, M. Rasel Mahmud