AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
- ID: f66a3ab6
- 原文链接: https://arxiv.org/abs/2608.13492
- PDF: https://arxiv.org/pdf/2608.13492v1
- 作者: AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
- 日期: 2026-08-13
- 更新: 2026-08-13
- 分类: models
- arXiv 分类: cs.AI
- 来源类型: paper
- 标签: world-model, long-horizon, spatial-memory, interactive, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-17T12:35:00Z
中文导读
AlayaWorld v1.1 技术报告聚焦交互式长时程世界模型的条件信号重构。设计原则只有一条:条件信号要在潜表示和时间结构上尽可能接近生成内容。两大改动:用流式 3D 点云缓存渲染器替换基于深度 warp 的空间记忆;把条件管线重设计为与生成视频同构的 causal-VAE 潜空间编码、时间统计一致。共六项修改,含运动感知潜条件、空间记忆因果连续编码、硬丢弃记忆 token(而非置零)、训练推理统一 VAE 协议、移除相机 AdaLN 分支。骨干分块自回归生成与训练数据均沿用上版。收录理由:世界模型 spatial memory 与 conditioning 设计取舍的一手实操记录。
为什么值得关注
长时程交互式世界模型的核心难题之一是"记忆"如何作为条件回流进生成管线。AlayaWorld v1.1 给出了一份少见的工程复盘:骨干、生成分块、训练数据全部不动,只重构条件信号的表示与注入方式,并明确列出六项修改的取舍逻辑(如硬丢弃 memory token 优于置零、相机控制完全交给重渲染空间条件)。对做 world model、视频生成条件设计、spatial memory 的读者,这是少有的"只动 conditioning"的消融式技术报告。
关键信息
- 论文标题:AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
- 作者:AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, et al.(按名首字母与贡献排序)
- arXiv:https://arxiv.org/abs/2608.13492
- 发布时间:2026-08-13
- arXiv 分类:cs.AI
- 关联标签:world-model / long-horizon / spatial-memory / interactive
English Abstract
This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two major changes. First, we replace the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer. Second, we redesign the conditioning pipeline so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.
English Summary
AlayaWorld v1.1 substantially revises how conditioning signals are represented in an interactive long-horizon world model, guided by one principle: conditioning should match the generated content as closely as possible in latent representation and temporal structure. Two major changes: depth-warping-based spatial memory is replaced by a streaming 3D point-cache renderer, and the conditioning pipeline is redesigned so visual conditions are encoded in the same causal-VAE latent space with consistent temporal statistics. Six concrete modifications include motion-aware latent conditioning, causally encoded re-rendered spatial memory as a continuous sequence, hard memory dropout that removes memory tokens rather than zeroing them, and unified VAE protocols across training and inference. Backbone, chunk-wise autoregressive generation, and training data are unchanged from the previous release.
Obsidian Notes
- 内容由
opencli arxiv paper 2608.13492拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。