Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li
Abstract
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
Create a lesson
Related papers
Moore, Escher, Penrose: A Conformal Golden Braid
Sophia Feldman, Assaf Shocher
Sphere Encoder 2
Kaiyu Yue, Sean McLeish, Ruchit Rawal et al.
One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars
Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
ROWBench: Do Video Models Render What the Program Specifies?
Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai et al.
Embedding Prediction Helps Image Generation
Sihan Xu, Ji Xie, Zilin Wang et al.
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.