ViSAR: Training-Free Adaptive-k Retrieval for Visual Document Question Answering
Adrien Mialland, Marc Plantevit, Julien Gallois, Céline Robardet
Abstract
Document Visual Question Answering (DocVQA) often leverages Retrieval-Augmented Generation (RAG), where late-interaction encoders are commonly used to identify document pages relevant to a user query, before answer generation by a Large Vision-Language Model (LVLM). Existing approaches typically retrieve a fixed top-k number of pages regardless of query complexity, which increases LVLM latency and may degrade answer accuracy. We introduce ViSAR (Visual Semantic Activation Retrieval), a training-free adaptive-k retrieval method for late-interaction visual document retrieval. ViSAR operates directly in the embedding space to construct a query-conditioned page-level similarity matrix that highlights query-relevant semantics and dynamically determines the number of pages to retrieve. Across multiple encoders and LVLMs, ViSAR retrieves compact, query-adapted page sets that reduce RAG latency by up to 58.7\%, while maintaining or improving answer accuracy compared with fixed top-k and adaptive retrieval heuristics. Furthermore, we show that the similarity matrix structure correlates with answer accuracy, suggesting future directions for retrieval quality-aware document understanding.
Create a lesson
Related papers
Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection
Max Nelson, Hanoz Bhathena, Aviral Joshi et al.
Recommender System as Slow and Fast Thinkers
Zichen Yuan, Xiaoxuan Dong, Linkun Dai et al.
Training seeds and model-selection stability in recommender-system evaluation
Juan Manuel Rodriguez, Oleg Lesota, Antonela Tommasel
Adaptive Test-Time Inference for Text2Cypher with Trace Budgeting and Selective Refinement
Makbule Gulcin Ozsoy
Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization
Bing Zheng, Zongyao Zhao, Wenming Yang
Genuine Information Needs of Social Scientists Looking for Data
Andrea Papenmeier, Thomas Krämer, Tanja Friedrich et al.