Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation
Ruofan Hu, Shengyang Xu, Minjie Hong, Xiaoda Yang, Sashuai Zhou, Ke Lei, Tao Jin, Zhou Zhao
Abstract
Real-world knowledge resides in multimodal documents, necessitating retrieval-augmented generation (RAG) for accurate question answering. However, existing multimodal RAG models are primarily designed for single-image or closed-document settings and exhibit limited accuracy in realistic multi-image scenarios. Moreover, processing numerous retrieved images incurs substantial computational overhead from irrelevant visual tokens. To address these challenges, we introduce DocLongRAG, a large-scale dataset of 343K question--answer pairs, each associated with an average of 37.4 retrieved images to reflect authentic RAG workflows. Building on this dataset, we propose Doc-REFRAG, a question-guided framework that compresses visual tokens into coarse chunks and selectively expands question-relevant ones via a lightweight RL-based selector. Experiments on six benchmarks show that Doc-REFRAG outperforms eleven strong baselines, achieving state-of-the-art accuracy with significantly lower inference latency. Our resources are available at https://github.com/Collab-Gen/Doc-REFRAG.
Create a lesson
Related papers
Closed Forms and Synthetic Twins: Predicting Approximate Nearest Neighbor Recall from Embedding Statistics
Shmuel Herman
MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval
Rohan Pandey, Sunjae Kwon, Hong Yu
Two-Sided State-Space Models for Sequential Recommendation with Non-Random Multimodal Review Feedback
Ziwen Pan, Zihan Liang, Ruoxuan Xiong
MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval
Seokwon Song, Sohyeon Kim, Gunhee Kim
Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval
Shaowei Wei, Chong Huang, Songtao Fang et al.
Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster
Songtao Fang, Zihao Xu, Shaowei Wei et al.