基础设施 4.0 · 优秀 2026-07-21 · 文章

NVIDIA Vera Rubin NVL72 on CoreWeave: 10x More Tokens Per Megawatt Than Blackwell

CoreWeave 发布 Vera Rubin NVL72 首个实测硅片性能:DeepSeek R1 工作负载matched interactivity target(TPS/user 对齐)条件下,相对 GB200 NVL72 生成 10x tokens-per-second per megawatt优化栈全开:大规模专家并行NVFP4multi-token predictionprefill/decode 分离(TensorRT-LLM + Dynamo)落地读法:同功率预算跑 10 倍 reasoning 流量或同流量省一个数量级功率,agentic 多步负载受益最直接不可外推边界:10x 是 tok/s/MW @ matched TPS/user,不是通用 TCO 常数;厂商自测无第三方复测;缺 TPS/user 目标的机架对比不要抄这个数...

打开原文回到归档

NVIDIA Vera Rubin NVL72 on CoreWeave: 10x More Tokens Per Megawatt Than Blackwell

Source: https://www.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell
Author: CoreWeave (Harsh Singh Banwait)
Published: 2026-07-21
Type: vendor blog / first measured silicon data
X post: https://x.com/i/status/2088375223724732608

中文导读

CoreWeave 发布 Vera Rubin NVL72 的首个实测硅片性能数据(2026-06 完成行业首次 bring-up 之后)。口径必须完整引用:DeepSeek R1 工作负载,在 matched interactivity target(TPS/user 对齐) 条件下,Vera Rubin NVL72 相对 GB200 NVL72 生成 10x tokens-per-second per megawatt。

优化栈全开:大规模专家并行、NVFP4 精度、multi-token prediction、prefill/decode 分离(TensorRT-LLM + Dynamo)。厂商给的落地读法:同功率预算跑 10 倍 reasoning 流量,或同流量省一个数量级功率 → 更低 cost-per-token,agentic 负载(多步工具调用+长推理)受益最直接。硬件侧:72 Rubin GPU + 36 Vera CPU、260 TB/s NVLink 6 全互联。

不可外推边界(引用时必须带):10x 是 tok/s/MW @ matched TPS/user,不是通用 TCO 常数;缺 TPS/user 目标的对比不要抄这个数;厂商自测、无第三方复测;CoreWeave 自己也写明这是起点,GB200 当年上线后数字随调优持续上涨。

Key Takeaways

  • 首个 Rubin NVL72 实测:R1 + matched interactivity 下 tok/s/MW ≈ 10x GB200 NVL72。
  • reasoning/agent 机架对比必须带交互性轴(TPS/user),只报吞吐会被 interactivity 目标暗改。
  • 引用模板:模型(R1)+ 指标(tok/s/MW)+ 匹配条件(TPS/user)+ 优化栈(TRT-LLM+Dynamo 全开)。
  • 功率测量点(PDU vs 板级)未在文中完整表格化,属待披露项。

Intake rationale

  • Category: infra | Quality score: 4/5
  • Rubin 能效的第一个实测锚点;对推理 infra/TCO 讨论是必引数据,且"matched interactivity"方法论比数字本身更有传播价值。

Grounding

opencli web read 原文全文(2026-08-16);数字与条件均出自原文;发布日期 2026-07-21 与作者经本地调研笔记(调研/2026-08-16-调研-Rubin-NVL72-tokMW与matched-interactivity.md)交叉核对。