Atlas: A World Model for Spatial Intelligence
- ID: e6922ed9
- 原文链接: https://www.worldlabs.ai/blog/atlas
- 作者: World Labs
- 日期: 2026-09-01
- 分类: models
- 来源类型: article
- 标签: world-model, spatial-intelligence, multimodal, 3d
- 质量评分: 4/5
- 抓取时间: 2026-09-02T12:28:56+00:00
中文导读
World Labs 发布 Atlas:一个从零预训练的多模态自回归扩散 Transformer,原生支持文本图像视频与 3D 输入它在所有见过的输入上保持 3D 一致性,并能在上下文之外继续想象;覆盖相机可控生成场景重建与仿真三类任务,性能随训练算力扩展而提升
为什么值得关注
World Labs 把 world model 和空间智能绑回一张统一的底座:Atlas 是一个从零预训练的 omni model,文本/图像/视频/3D 在同一个空间上下文里编码,输出仍然是多模态的,但每张图都被钉在一个 3D 位置上。它的能力面覆盖三件事:相机可控生成、最少一张图就能拿下的稀疏视角 3D 重建,以及 Real-to-Sim 风格的时空仿真;benchmark 上 Atlas 在相机控制生成上超越近期视频模型,在稀疏输入 3D 重建上也超越当前最强的开源专用方法,且团队明确表示性能随训练算力继续 scale。下游会接入 Marble 等产品,正在做早期合作伙伴接入。对关注生成式 3D/spatial intelligence 与具身仿真的人来说,这是最近几个月最值得跟进的体系级发布之一。
关键信息
- 文章标题: Atlas: A World Model for Spatial Intelligence
- 发布方: World Labs
- 原文链接: https://www.worldlabs.ai/blog/atlas
- 发布时间: 2026-09-01
- 核心能力: 相机可控生成、稀疏视角 3D 重建、时空仿真、图像生成
- 模型形态: multimodal autoregressive diffusion transformer(native 文本/图像/视频/3D 输入)
- 关联标签: world-model, spatial-intelligence, multimodal, 3d
English Summary
World Labs introduces Atlas, a multimodal autoregressive diffusion transformer pretrained from scratch on text, images, video and 3D. Atlas maintains 3D consistency across the inputs it has seen, can extrapolate beyond the visible context, and supports camera-controlled generation, scene reconstruction, and simulation. Performance is reported to keep scaling with training compute.
Obsidian Notes
- 内容由
opencli web read拉取 World Labs 博客原文生成;正文保存在web-articles/Atlas_A_World_Model_for_Spatial_Intelligence/目录作为原始素材。 - 中文导读与价值判断均锚定在条目已有摘要、博客原文小节标题与官方描述上;具体 benchmark 数字与架构细节以官方博客为准,本文件不做额外外推。