模型与实验室 4.0 · 优秀 2026-10-01 · 文章

Understanding the AI That Drives Robots

Construction Physics 的人形机器人 VLA(vision-language-action model)长文入门,从线性代数attentiontransformer 一路讲到 Physical Intelligence 的开源 0.5:VLM 主干用 PaliGemma(SigLIP 400M 图像编码器 + Gemma 2B),并行一个 18 层300M 参数的 action transformer,输入 50 个随机动作 embedding,每层 attention 都回头读 VLM 对应块的图像/文本/机器人状态向量,迭代 10 次把噪声磨成 50 步关节角/抓手目标...

打开原文回到归档

Understanding the AI That Drives Robots

Intake entry · 2026-10-04 · awesome-ai-field-notes

中文摘要

Construction Physics 的人形机器人 VLA(vision-language-action model)长文入门,从线性代数、attention、transformer 一路讲到 Physical Intelligence 的开源 π0.5:VLM 主干用 PaliGemma(SigLIP 400M 图像编码器 + Gemma 2B),并行一个 18 层、300M 参数的 action transformer,输入 50 个随机动作 embedding,每层 attention 都回头读 VLM 对应块的图像/文本/机器人状态向量,迭代 10 次把噪声磨成 50 步关节角/抓手目标;一次输出 50 个动作(action chunking)避免每动一步重跑整个模型,动作用 FAST 方案 token 化,最后交给传统控制器转成电机扭矩。作者判断:VLA 与通用 LLM 共享的结构多到意外(直接拿 LLM 当主干、transformer 复用来生成动作),但通用 AI 的快速能力跃迁不一定复制到机器人上,也不排除可能。文首盘点 Figure 10 亿、Apptronik 5.2 亿、Unitree 上市募 9 亿美元等融资。

关键要点

  • VLA = 文本+图像+机器人状态进 transformer,输出动作序列而非文本;Figure/Unitree/Physical Intelligence/英伟达都走这条路线。
  • π0.5 具体构造:PaliGemma 主干 + 18 层 300M action transformer + 50 动作并行 + 10 次迭代 + action chunking + FAST 动作 token 化。
  • 机器人进步主要来自 AI 而非硬件;但 demo 要打折扣看(Figure 分拣包裹、Physical Intelligence 咖啡/折盒)。
  • 结构上 VLA 直接复用文本训练的 LLM 当主干,但通用 LLM 的能力跃迁速度未必复制到机器人。

English Summary

A long-form Construction Physics explainer of vision-language-action models, from linear algebra and attention to Physical Intelligence's open-weight pi0.5: a PaliGemma VLM backbone (SigLIP 400M image encoder + Gemma 2B) runs alongside an 18-block, 300M-parameter action transformer fed 50 random action embeddings; at every attention layer each action vector reads the VLM's image/text/state vectors, and after 10 iterations the noise is refined into 50 joint-angle/gripper targets, output as one chunk (action chunking) so the model is not rerun per step. Actions are tokenized with FAST and handed to a traditional controller. Potter notes VLAs share surprising structure with LLMs but general-AI capability jumps may not transfer to robots.

One-liner

Construction Physics 的人形机器人 VLA(vision-language-action model)长文入门,从线性代数attentiontransformer 一路讲到 Physical Intelligence 的开源 0.

原文摘录

A VLA takes text, images, and robot information as an input and gives a series of robot actions as an output. ... After 10 iterations, the output of the action transformer is, hopefully, a useful sequence of 50 actions.
注:本文件为 daily-intake-evening cron 入库;opencli web read 抓取全文后提取,原文全文见抓取缓存。