Abstract
Understanding information from a collection of multiple documents, particularly those with visually rich elements, is important for document-grounded question answering. This paper introduces VisDoMBench, the first comprehensive benchmark designed to evaluate QA systems in multi-document settings with rich multimodal content, including tables, charts, and presentation slides. We propose VisDoMRAG, a novel multimodal Retrieval Augmented Generation (RAG) approach that simultaneously utilizes visual and textual RAG, thereby combining robust visual retrieval capabilities with sophisticated linguistic reasoning. VisDoMRAG employs a multi-step reasoning process encompassing evidence curation and chain-of-thought reasoning for concurrent textual and visual RAG pipelines. A key novelty of VisDoMRAG is its consistency-constrained modality fusion mechanism, which aligns the reasoning processes across modalities at inference time to produce a coherent final answer. This leads to enhanced accuracy in scenarios where critical information is distributed across modalities and improved answer verifiability through implicit context attribution. Through extensive experiments involving open-source and proprietary large language models, we benchmark state-of-the-art document QA methods on VisDoMBench. Extensive results show that VisDoMRAG outperforms unimodal and long-context LLM baselines for end-to-end multimodal document QA by 12-20%.
The problem
In practice, a question is rarely asked of a single PDF. Analysts, scientists, and lawyers ask questions over a folder of documents, and the system first has to find the one that answers the question, then find the page, table, or chart inside it. That is a needle-in-a-haystack problem, and the needle is often not text: a number in a table, a trend in a plot, or a bullet on a slide.
Existing multi-document QA benchmarks are almost entirely textual, and existing multimodal document QA work is single-document. There was no benchmark for the combination, and no clear answer to a simple engineering question: when documents mix text and visuals, should you retrieve page images, extracted text, or both, and how should the two be combined?
Approach
VisDoMBench re-purposes five document QA datasets that satisfy three criteria: visually rich content, publicly available source documents, and grounded evidence. The splits are PaperTab and FetaTab (tables, via the UDA benchmark from QASPER and FeTaQA), SciGraphQA and SPIQA (charts and tables from scientific papers), and SlideVQA (multi-hop questions over slide decks). We de-duplicate across splits, filter trivial questions, and augment every question with distractor documents so that each query spans roughly 50 to 200 pages in total. Ambiguous questions are rewritten with GPT-4o and reviewed by a human annotator so that exactly one document answers each question. The result is 2,271 questions over 1,277 documents, with an average of 8.4 documents and about 129 pages per query.
VisDoMRAG answers a query with two parallel, evidence-driven unimodal RAG pipelines followed by a fusion step.
- Textual RAG. OCR the documents, chunk the text (3,000 characters with overlap), index it with a text embedding model, and retrieve the top chunks for the query.
- Visual RAG. Index every page as an image with a late-interaction visual retriever (ColPali or ColQwen2) and retrieve the top pages, which go to a multimodal LLM as images.
- Three-step prompting in each branch. Evidence Curation asks the LLM to extract and verbalize the relevant paragraphs, tables, or figure details from the retrieved context; Chain-of-Thought Reasoning links that evidence into a step-by-step argument; Answer Generation produces a response in the format the question type calls for.
- Modality Fusion. A final LLM call receives the curated evidence, reasoning chains, and answers from both branches and checks them for consistency, reconciling contradictions and filling reasoning gaps before producing the final answer. This is a late-fusion design, inspired by self-consistency over chains of thought.
Example
The paper’s opening example (Figure 1, shown above under “The problem”) is a question whose answer sits in one table inside one of the documents in a collection.
Question over a document collection
Grounding context (one table, in one of the documents)
| Length | [0, 30) | [30, 60) | [60, 90) | [90, ∞) |
|---|---|---|---|---|
| #Pair | 253578 | 207772 | 33618 | 5032 |
| LSTM | 0.707 | 0.748 | 0.732 | 0.718 |
| MV-LSTM | 0.726 | 0.752 | 0.725 | 0.694 |
| KEHNN | 0.724 | 0.774 | 0.785 | 0.791 |
Answer
The second example is the paper’s qualitative comparison on PaperTab (Figure 5). The two unimodal pipelines each latch onto a different wrong number from the same table; the fused reasoning gets the sum right.
Query
Ground-truth evidence and answer
Visual RAG
Textual RAG
VisDoMRAG
Results
| Method | LLM | PaperTab | FetaTab | SciGraphQA | SPIQA | SlideVQA | Average |
|---|---|---|---|---|---|---|---|
| Long Context | Qwen2-VL | 8.23 | 23.10 | 16.74 | 9.93 | 2.46 | 12.09 |
| Long Context | GPT-4o | 28.37 | 60.03 | 24.12 | 36.30 | 15.06 | 32.78 |
| Text RAG | GPT-4o | 37.34 | 60.82 | 29.74 | 42.80 | 15.97 | 37.33 |
| Visual RAG | GPT-4o | 42.01 | 61.89 | 31.12 | 43.28 | 66.82 | 49.02 |
| VisDoMRAG | Qwen2-VL | 29.89 | 59.24 | 27.98 | 42.80 | 39.77 | 39.94 |
| VisDoMRAG | Gemini 1.5 Flash | 39.66 | 60.89 | 25.82 | 41.03 | 52.74 | 44.03 |
| VisDoMRAG | GPT-4o | 44.11 | 63.28 | 31.36 | 44.09 | 67.22 | 50.01 |
Source: Table 3 of the paper. End-to-end QA accuracy (%) on the five VisDoMBench splits, judged against the ground-truth answer. Higher is better. The full table also reports Qwen2-VL and Gemini for every baseline.
- VisDoMRAG beats long-context, text-only RAG, and visual-only RAG for every LLM tested. Per-dataset gains over the baselines range from 2.1-21.6 points on PaperTab to 0.4-52.2 on SlideVQA.
- The gain is largest for the smallest open model: Qwen2-VL goes from 12.09 with long context to 39.94 with VisDoMRAG.
- Long-context models degrade as the page count per query grows (Figure 4); VisDoMRAG stays flat because retrieval bounds the context.
- Ablations with GPT-4o (Table 5): removing evidence curation, chain-of-thought, and consistency prompting drops VisDoMRAG from 50.01 to 45.98; early fusion (appending OCR text of the visually retrieved pages to the visual context) scores 43.63, below late fusion.
Which retriever finds the right document?
| Retriever | Modality | PaperTab | SlideVQA | Average |
|---|---|---|---|---|
| BM25 | text | 65.51 | 98.55 | 81.80 |
| MiniLM | text | 65.51 | 0.73 | 61.56 |
| MPNet | text | 90.18 | 0.73 | 73.57 |
| BGE-1.5 | text | 96.81 | 81.85 | 92.40 |
| ColPali | visual | 96.93 | 97.64 | 96.15 |
| ColQwen2 | visual | 97.61 | 97.82 | 96.94 |
Source: Table 4 of the paper. Share of queries (%) for which the ground-truth document supplies the majority of the top-5 retrieved pages or chunks. Higher is better. Dense text retrievers nearly fail on SlideVQA because slides carry sparse keyword text, while BM25 matches those keywords directly.
VisDoMBench statistics
| Split | Domain | Content | Queries | Docs | Avg. docs / query | Avg. pages / query |
|---|---|---|---|---|---|---|
| PaperTab | Scientific papers | Tables, text | 377 | 297 | 10.82 | 113.10 |
| FetaTab | Wikipedia | Tables | 350 | 300 | 7.77 | 124.33 |
| SciGraphQA | Scientific papers | Charts | 407 | 319 | 5.91 | 129.71 |
| SPIQA | Scientific papers | Tables, charts | 586 | 117 | 9.51 | 135.58 |
| SlideVQA | Presentation decks | Slides | 551 | 244 | 6.99 | 139.71 |
| VisDoMBench | Combined | Tables, charts, slides, text | 2,271 | 1,277 | 8.36 | 128.69 |
Source: Table 2 of the paper (domain labels for PaperTab and FetaTab follow the source datasets, QASPER and FeTaQA). Each query is paired with distractor documents so that the collection spans roughly 50-200 pages.
Resources
- Paper on arXiv and the ACL Anthology page
- Code and VisDoMBench on GitHub: the five data splits plus
visdomrag.py, the VisDoMRAG implementation (ColPali/ColQwen visual retrieval; BM25, MiniLM, MPNet, BGE text retrieval; GPT-4, Gemini, Qwen backends) - Hugging Face Papers page
- Adobe Research publication page
- Explainer post on this site, and all publications
Quick start
From the repository README:
from visdomrag import VisDoMRAG
config = {
"data_dir": "./path/to/dataset",
"output_dir": "./results",
"llm_model": "gpt4", # "gpt4", "gemini", "qwen"
"vision_retriever": "colpali", # "colpali", "colqwen"
"text_retriever": "bm25", # "bm25", "minilm", "mpnet", "bge"
"api_keys": {"openai": "your-openai-key", "gemini": "your-gemini-key"},
}
pipeline = VisDoMRAG(config)
pipeline.run() # process all queries
# or: pipeline.process_query(query_id)
