Agent 与自动化 4.0 · 优秀 2026-09-09 · 论文

Show-Harness: Just a VLM Agent Can Play Robots

怎么让 VLM 的世界知识变成机器人控制?Show-Harness 给的答案是"具身 harness":暴露离散的语义动作单元让 VLM 推理,由具体本体(embodiment)的解释器确定性地落到本地机器人动作,细粒度物理决策仍由 VLM 直接负责 同一接口验证两件事:闭源 frontier VLM 直接零样本控制机器人;小规模开源 VLM 只用几个 GPU 小时微调即可低成本部署 还做了 GUMI(GUI Manipulation Interface):把同一语义动作空间扩展到 GUI 示教采集,人和智能体都能"玩"机器人,不需要专业遥操作硬件 实验显示 Show-Harness 加持的 VLM 智能体跨任务/本体/环境泛化稳健,超过代表性 agentic 和 VLA 范式...

打开原文回到归档

Show-Harness: Just a VLM Agent Can Play Robots

  • ID: 237d5343
  • 原文链接: https://arxiv.org/abs/2609.10522
  • PDF: https://arxiv.org/pdf/2609.10522v1
  • 作者: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
  • 日期: 2026-09-09
  • 更新: 2026-09-09
  • 分类: cs.RO, cs.AI, cs.CV, cs.MM
  • 来源类型: paper
  • 标签: vlm, robotics, embodied-agents, harness, teleoperation, field-note
  • 质量评分: 4/5
  • 抓取时间: 2026-09-11T12:23:09+00:00

中文导读

怎么让 VLM 的世界知识变成机器人控制?Show-Harness 给的答案是"具身 harness":暴露离散的语义动作单元让 VLM 推理,由具体本体(embodiment)的解释器确定性地落到本地机器人动作,细粒度物理决策仍由 VLM 直接负责 同一接口验证两件事:闭源 frontier VLM 直接零样本控制机器人;小规模开源 VLM 只用几个 GPU 小时微调即可低成本部署 还做了 GUMI(GUI Manipulation Interface):把同一语义动作空间扩展到 GUI 示教采集,人和智能体都能"玩"机器人,不需要专业遥操作硬件 实验显示 Show-Harness 加持的 VLM 智能体跨任务/本体/环境泛化稳健,超过代表性 agentic 和 VLA 范式;结论是"对的接口"能从基础 VLM 里解锁大量具身能力,不需要加容量或昂贵的本体预训练

为什么值得关注

VLM 控机器人的接口设计:语义动作单元给 VLM 推理,本体解释器确定性落地;闭源 frontier 模型零样本可用,小开源模型几个 GPU 小时微调即部署,GUMI 还免遥操作硬件

条目锚定 arXiv 2609.10522(2026-09-09 提交,分类 cs.RO, cs.AI, cs.CV, cs.MM),摘要自述贡献为上述机制与结论;详细信息以论文原文为准。

关键信息

  • 论文标题: Show-Harness: Just a VLM Agent Can Play Robots
  • 作者: Yanzhe Chen, Zechen Bai, Zhijun Cao, Wenzheng Zeng, Kevin Qinghong Lin, Yiqi Lin, Guoqiang Liang, Kevin Yuchen Ma, Qiming Huang, Mike Zheng Shou
  • arXiv: https://arxiv.org/abs/2609.10522
  • 发布时间: 2026-09-09
  • arXiv 分类: cs.RO, cs.AI, cs.CV, cs.MM
  • 关联标签: vlm, robotics, embodied-agents, harness, teleoperation, field-note

English Abstract

Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness demonstrates the feasibility of (1) directly unlocking closed-source frontier VLMs for zero-shot robot control, and (2) adapting small-scale open-source VLMs for low-cost deployment with just a few GPU-hours of fine-tuning. We further develop GUMI (GUI Manipulation Interface), which extends the same semantic action space to GUI-based demonstration collection, allowing humans and agents to "play" robots across embodiments without specialized teleoperation hardware. Extensive experiments show that Show-Harness-equipped VLM agents generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms. These results suggest that the right interface can unlock substantial embodied capability from foundation VLMs, without requiring additional model capacity or costly embodiment-specific pretraining.

English Summary

Show-Harness is an Embodied Harness enabling VLMs to "play" robots through a compact semantic interface: discrete semantic action units that VLMs reason over are deterministically grounded into local robot actions by embodiment-specific interpreters, while the VLM stays responsible for fine-grained physical decisions. The same interface unlocks closed-source frontier VLMs for zero-shot robot control and adapts small open-source VLMs with just a few GPU-hours of fine-tuning. GUMI extends the semantic action space to GUI-based demonstration collection without specialized teleoperation hardware. VLM agents equipped with Show-Harness generalize robustly across tasks, embodiments, and environments, outperforming representative agentic and VLA paradigms.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。