VISTA: A Visual Harness for Reasoning in an Interactive World
Content-fetcher entry · 2026-10-04 · awesome-ai-field-notes
- URL: https://arxiv.org/abs/2610.02200
- PDF: https://arxiv.org/pdf/2610.02200
- Source: arxiv · Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, Kaiming He
- Original Date: Thu, 1 Oct 2026 17:59:45 UTC (4,452 KB)
- Added: 2026-10-04
- Category: agents
- Tags: agents, multimodal, visual-memory, harness, arc-agi, benchmark, interactive-environments
- Quality Score: 5
中文摘要
VISTA:给通用多模态模型装上长时程视觉的外挂框架模型直接通过视觉观察感知交互环境,并维护无损视觉记忆(以原始形态保存历史观察),推理时可主动检索并重组自己的视觉输入在 ARC-AGI-3 上,VISTA 把 Claude Opus 5.0 的 Relative Human Action Efficiency 从 40.68 提到满分 100.00,25 个公开游戏全部通关,动作用量比首次游玩的人类少 57.4%;同一设计在另外三个视觉游戏/谜题基准上也大幅超过同底模配简单 harness 的基线作者含 Kaiming HeCathy Wu;结论指向:多模态模型本身推理能力不弱,瓶颈在外部 harness2026-10-01,cs.AI/cs.CV
English Abstract
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
One-liner
VISTA:给通用多模态模型装上长时程视觉的外挂框架模型直接通过视觉观察感知交互环境,并维护无损视觉记忆(以原始形态保存历史观察),推理时可主动检索并重组自己的视觉输入在 ARC-AGI-3 上,VISTA 把 Claude Opus 5.
注:本文件为 content-fetcher cron 回填;opencli arxiv 拉取失败(COMMAND_EXEC),改以 opencli web read 抓取 arXiv abs 页面,仅依据可见的标题/作者/提交时间/分类/摘要生成,未补充摘要之外的实验细节。