OpenAI 与 Broadcom 联合发布 LLM 优化推理芯片 Jalapeño
- ID: 1a015f5d
- 原文链接: https://openai.com/index/openai-broadcom-jalapeno-inference-chip
- 作者: OpenAI / Broadcom
- 日期: 2026-06-24
- 分类: infra
- 标签: OpenAI, Broadcom, 推理芯片, Jalapeño
- 质量评分: 4/5
- 抓取时间: 2026-06-29T20:26:43
中文翻译
概览
OpenAI 与 Broadcom(NASDAQ: AVGO)今日联合发布了 Jalapeño —— OpenAI 的首款"智能处理器",专为 OpenAI 对 LLM 推理未来的愿景而设计,也是双方共同打造的多代际计算平台中的首款 AI 加速器,旨在让先进 AI 更快、更可靠、更易于为更多人所用。
Jalapeño 由 Broadcom 总裁兼 CEO Hock Tan、总裁 Charlie Kawwas 亲自交付给 OpenAI CEO Sam Altman 与总裁 Greg Brockman,标志着 OpenAI 全栈模型与产品战略迈出了重要一步。
设计要点
- 早期测试显示,第一代加速器将在性能功耗比上显著优于当前业界最先进水平
- 从零开始为整个行业当前与未来的 LLM 而设计
- 从设计到生产仅用九个月,并由 OpenAI 模型加速
- 扩展了 OpenAI 的全栈平台:从产品到模型,再到芯片
- 将与数据中心合作伙伴以吉瓦级规模部署,覆盖多代际产品
OpenAI 从头设计了这款芯片,基于其对 LLM 底层的深度理解,并结合其模型、内核、服务系统和产品需求的路线图。Broadcom 与 Celestica 作为合作伙伴,帮助实现该平台的工业化:芯片实现、板卡、机架系统集成、高性能网络以及可扩展的生产系统。Jalapeño 设计上具备灵活性,可在 OpenAI 对当前和未来 AI 模型推理需求的洞察指导下,与所有 LLM 协同工作。Jalapeño 芯片的工程样片已在实验室中以生产目标频率和功耗运行 ML 工作负载,包括 GPT‑5.3‑Codex‑Spark。
关键引述
Greg Brockman(OpenAI 总裁兼联合创始人):"世界正在迈向计算驱动的经济。Jalapeño 是我们长期全栈基础设施战略的一部分,旨在让算力更加充裕,从而使 AI 对个人与企业而言更快、更可靠、更可负担,并可用于解决更重要的问题。通过自行设计更多层栈,我们可以更高效地服务更多智能,并将先进 AI 推向更广泛的访问。"
Richard Ho(OpenAI 硬件项目负责人):"Jalapeño 是利用我们与 OpenAI 研究人员紧密协作所获得的深入洞察,从零开始为 LLM 推理而设计的。我们围绕对前沿 AI 模型最重要的内核、内存移动、网络和服务模式,对架构进行了优化。基于早期测试,Jalapeño 将能够在接近硬件理论极限的情况下高效执行我们最重要的工作负载。"
Hock Tan(Broadcom 总裁兼 CEO):"我们与 OpenAI 的合作代表了对未来十年 AI 所需物理基础设施扩展的根本承诺。这仅仅是个开端,是一个多代际路线图的起点。通过与 OpenAI 直接共同开发我们的业界领先硅片,我们正与 Microsoft 等合作伙伴一道,自 2026 年起实现吉瓦级数据中心的部署。"
为 LLM 设计的最佳推理平台
Jalapeño 是面向现代 LLM 推理的白板设计,而非从早期 AI 工作负载改造而来的通用加速器。它的设计基于 OpenAI 在 ChatGPT、Codex、API 以及未来智能体产品中每天运行系统的经验,同时也面向整个行业当前与未来的 LLM。目标是将当今领先 AI 加速器的强大能力与吞吐量,与最快专用推理系统的延迟相结合,使 Jalapeño 非常适合大规模交互式 LLM 产品。
这就是全栈优势。OpenAI 不仅在开发前沿模型或在其之上构建产品,更在设计其下的基础设施:芯片架构、内核、内存系统、网络、调度、部署系统和产品体验。由于 OpenAI 横跨整个堆栈,每一层都可以围绕同一个目标进行优化:让模型对用户而言更快、更可靠、更便宜。
Jalapeño 强化了 OpenAI 进步背后的飞轮:更好的基础设施驱动计算效率;更高的计算效率支持更好的训练和服务,最终赋能更强大的 AI 模型;更好的模型成为面向个人、开发者和企业的更优产品;更好的产品带来更多使用、更多客户和更多收入,进而使 OpenAI 能够对下一代基础设施进行再投资。随着时间推移,这一循环让智能对每个人而言都更强大、更可靠、更便宜。
九个月流片,由 OpenAI 模型加速
Jalapeño 从最初设计到制造流片仅用九个月共同开发完成,这一定制 AI 加速器项目据信是高性能先进半导体领域有史以来最快的 ASIC 开发周期。这一速度反映了与 OpenAI 工程团队的深度软硬件协同开发、Broadcom 的硅片实现专长,以及 OpenAI 模型在设计和优化部分环节中的运用。
服务用户的同一批模型正在帮助改进运行未来模型的基础设施。如果 AI 能帮助工程师更快地设计出更好的芯片,它就可以降低整个行业的计算成本,并帮助普及对先进 AI 的访问。
与合作伙伴共建多代际平台
Jalapeño 是多代际计算平台的第一步,该平台计划于 2026 年底首次部署,并在未来几年扩展,将 OpenAI 设计的加速器与 Broadcom 的硅片实现、网络和连接技术以及 Celestica 的板卡、机架和系统专长相结合。
让先进 AI 更广泛可用
这项工作的意义很简单:推理是 AI 触达用户的环节。每一次成本、速度和可靠性的改进,都可能体现为更快的 ChatGPT 答案、能承担更多步骤而等待更少的 Codex 任务、更便宜的 API 产品,或在高需求时更可靠的访问。
English Original
OpenAI and Broadcom unveil LLM-optimized inference chip
*June 24, 2026 — OpenAI*
- Early testing shows the first-generation accelerator will deliver performance per watt substantially better than current state-of-the-art
- Built from the ground up for current and future LLMs across the industry
- Developed from design to production in nine months, accelerated by OpenAI's models
- Expands OpenAI's full-stack platform, from products to models and now to chips
- To be deployed at gigawatt scale with data center partners, over multiple generations
OpenAI and Broadcom (NASDAQ: AVGO) today unveiled Jalapeño, OpenAI's first Intelligence Processor — an accelerator architected around OpenAI's vision for the future of LLM inference, and the first AI accelerator in a multi-generation compute platform the companies are building together to make advanced AI faster, more reliable, and more accessible.
OpenAI designed the chip from scratch around its deep understanding of LLM fundamentals, informed by its roadmap of models, kernels, serving systems, and product needs, with partners Broadcom and Celestica helping industrialize the platform through chip implementation, board, rack system integration, high-performance networking, and scalable production systems. Engineering samples of Jalapeño are running ML workloads in the lab at production target frequency and power, including GPT‑5.3‑Codex‑Spark.
Designed to be the best inference platform for LLMs
Jalapeño is a blank-slate design for modern LLM inference, not a general-purpose accelerator adapted from earlier AI workloads. It is informed by the systems OpenAI runs every day across ChatGPT, Codex, the API, and future agentic products, while also being designed for current and future LLMs across the industry.
That is the full-stack advantage. OpenAI is not only developing frontier models or building products on top of them; it is designing the infrastructure underneath them: chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience.
Nine-month tape-out, accelerated by OpenAI models
Jalapeño was co-developed from initial design to manufacturing tape-out in just nine months, the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors. That speed reflects deep software-hardware co-development with OpenAI's engineering teams, Broadcom's silicon implementation expertise, and the use of OpenAI models to accelerate parts of the design and optimization process.
Building a multi-generation platform with partners
Jalapeño is the first step in a multi-generation compute platform designed for initial deployment by the end of 2026 and expanding in the years ahead, combining OpenAI-designed accelerators with Broadcom silicon implementation, networking, and connectivity technologies, and Celestica's board, rack, and system expertise.
Making advanced AI more broadly available
The point of this work is simple: inference is where AI reaches people. Every improvement in cost, speed, and reliability shows up as a faster ChatGPT answer, a Codex task that can take more steps with less waiting, an API product that is cheaper to build, or more dependable access when demand is high.