VistaHop: Benchmarking Long-Horizon Visual DeepSearch

A benchmark for repeated image inspection, visual-anchor grounding, and long-horizon evidence traversal across different visual regions.

Hang He1,2,3 Chuhuai Yue2 Chengqi Dong4,2 Chengcheng Wan1,3,† Ting Su1 Haiying Sun1
Jiajun Chai2 Xiaohan Wang2 Guojun Yin2,†

Corresponding authors

600tasks
600images
25scenarios
74.3%L3 tasks
14.92avg. evidence steps
01

Overview

Abstract

Visual DeepSearch tasks require multimodal large reasoning models to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps. However, existing benchmarks primarily evaluate single-step visual understanding or isolated visual-query response generation. They have limited difficulty, limited search horizons, and single-pass image inspection, and thus fail to evaluate models' ability to iteratively revisit visual evidence and reason across multiple steps.

We introduce VistaHop, a benchmark designed specifically to evaluate Visual DeepSearch. It evaluates repeated image inspection, visual-anchor grounding, and long-horizon evidence traversal across different visual regions. VistaHop comprises 600 images, 25 visual search scenarios, and 600 Visual DeepSearch tasks. We also propose VistaArena, a unified evaluation framework that supports tool-based interactions, including visual retrieval, image inspection, and evidence-grounded reasoning. Experiments show that even state-of-the-art MLRMs remain far from solving VistaHop, with the best-performing model, SenseNova-MARS-32B, achieving only 26.33% Pass@1. These findings highlight the importance of specialized benchmarks and improved agentic methods for Visual DeepSearch.

C1

VistaHop benchmark

Evaluates long-horizon evidence traversal, repeated image inspection, and evidence-grounded response generation through multi-chain Visual DeepSearch tasks.

C2

Scalable construction

Generates visually grounded, long-horizon tasks while controlling data quality and reducing text-only shortcuts.

C3

VistaArena evaluation

Shows that current models remain limited in long-horizon evidence traversal, visual evidence revisiting, and multi-anchor evidence fusion.

02

Task Synthesis

VistaHop Construction Pipeline

Seven-stage VistaHop benchmark construction pipeline

Figure 2 Overview of the VistaHop Construction Pipeline.

1–2Ground

Filter high-resolution images, detect entities, and enrich visual anchors.

3–4Connect

Build evidence chains and write indirect, evidence-grounded queries.

5–6Stress test

Remove shortcuts, verify clue necessity, and fuse multiple chains.

7Verify

Human reviewers check uniqueness, correctness, and readability.

03

Reviewer dataset

Explore all 600 VQA tasks

Every record pairs a web-optimized image with its task query and reference answer. Start with the featured cases, then filter the complete release by difficulty, category, or reasoning form.

Machine-readable release

Images + task queries + references

One self-contained JSONL includes all 600 records. Every line stores the image as Base64-encoded JPEG bytes together with the task query, reference, and complete task metadata.

600VQA records
154L2 tasks
446L3 tasks
138single-chain
462multi-chain

Featured examples

Five representative cases

Cross-domain examples selected for clear visual anchors, multi-entity reasoning, and natural task wording.

Full benchmark

Browse the complete release

Loading 600 records…
04

Evaluation

VistaArena Evaluation Framework

VistaArena overview showing the Search Agent, multimodal tools, iterative search trajectory, and Validation Agent decision flow

Figure 7 Overview of VistaArena. The Search Agent performs iterative visual inspection, retrieval, and evidence-grounded reasoning; the Validation Agent evaluates the resulting trajectory and final answer.

Text Search

Retrieve external facts and condense relevant webpage evidence.

Image Search

Use reverse-image evidence to resolve uncertain visual anchors.

Image Crop

Inspect local regions for small logos, text, objects, and symbols.

Validation Agent

Judge semantic correctness and evidence support with Pass@1.

05

Experiments

Overall Performance

26.33%best Pass@1

The best-performing model, SenseNova-MARS-32B, reaches only 26.33% Pass@1 under Search+Crop. The human baseline reaches 79.50%.

Search+Crop Pass@1 (%)

SenseNova-MARS-32B26.33
GPT-5.225.83
Claude Sonnet 4.523.68
Gemini 2.5 Pro21.83
Qwen3-VL-235B20.54
Qwen3-VL-30B18.34
MMSearch-R1-7B17.18

What makes it hard?

Difficulty distribution comparison showing that 74.3 percent of VistaHop tasks are L3, substantially more than in the other visual search benchmarks

VistaHop consists of L2 and L3 tasks, with 74.3% categorized as L3, placing substantially greater emphasis on long and compositional evidence chains.

Project reach

Visitor Statistics

A live geographic view of the VistaHop research community.

Powered by MapMyVisitors. Open the detailed view for real-time visitor statistics.

Task query

Reference