Agent 与自动化 4.0 · 优秀 2026-08-04 · 论文

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Video-DeepResearch 把多模态 deep research 从静态图像扩展到连续视频流,指出当前模型存在绕过视觉工具的 modality bias 和依赖参数记忆的 leakage;论文引入 Video-DRBench,用 dense spatiotemporal grounding 与 open-web exploration 同时测试视频研究 Agent

打开原文回到归档

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

  • ID: 166c7bcc
  • 原文链接: https://arxiv.org/abs/2608.03979
  • PDF: https://arxiv.org/pdf/2608.03979v1
  • 作者: Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao
  • 日期: 2026-08-04
  • 更新: 2026-08-04
  • 分类: agents
  • 来源类型: paper
  • 标签: multimodal-agents, deep-research, video, evaluation
  • 质量评分: 4/5
  • 抓取时间: 2026-08-06T04:19:06Z

中文导读

Video-DeepResearch 把多模态 deep research 从静态图像扩展到连续视频流,指出当前模型存在绕过视觉工具的 modality bias 和依赖参数记忆的 leakage;论文引入 Video-DRBench,用 dense spatiotemporal grounding 与 open-web exploration 同时测试视频研究 Agent

为什么值得关注

视频研究 Agent 的关键瓶颈是时空 grounding工具使用偏置与知识泄漏

Video-DeepResearch 把多模态 deep research 从静态图像扩展到连续视频流,指出当前模型存在绕过视觉工具的 modality bias 和依赖参数记忆的 leakage;论文引入 Video-DRBench,用 dense spatiotemporal grounding 与 open-web exploration 同时测试视频研究 Agent

关键信息

  • 论文标题:Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
  • 作者:Zhen Fang, Yu Zeng, Wenxuan Huang, Yiming Zhao, Shiting Huang, Tianfei Ren, Qi Lu, Qingnan Ren, Qisheng Su, Lionel Z. Wang, Qingyu Yin, Shuang Chen, Zehui Chen, Lin Chen, Zhenfei Yin, Yao Hu, Shaohui Lin, Wanli Ouyang, Shaosheng Cao, Feng Zhao
  • arXiv:https://arxiv.org/abs/2608.03979
  • 发布时间:2026-08-04
  • arXiv 分类:cs.CV, cs.AI
  • 关联标签:multimodal-agents, deep-research, video, evaluation

English Abstract

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

English Summary

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval....

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。