模型与实验室 4.0 · 优秀 2026-09-23 · 文章

Ember-1: a Kimi K3 reasoning model with 40% fewer tokens

Fireworks Research 9 月 23 日发布 Ember-1:在 Kimi K3 基础上训练的专用模型,号称比 K3 输出 token 减少约 40%,多项外部基准两个付费客户的线上 A/B 测试与自家编码/agent 工作负载上质量持平团队观察到 K3 在生成 token 上有超过 90% 实际用于内部推理,而多轮 agent 工作流会把历史推理完整回灌,导致上下文随轮次近似二次增长Fireworks 跑了 50+ 训练实验和 200+ 评估(在 Serverless Training 平台上完成),保留必要的反思性推理并裁掉冗余循环...

打开原文回到归档

Ember-1: a Kimi K3 reasoning model with 40% fewer tokens

原文链接: https://fireworks.ai/blog/ember-1
作者: Fireworks Research
发布时间: 2026-09-23
源: arXiv外部扫描 (2026-09-28)

摘要

Fireworks Research 9 月 23 日发布 Ember-1:在 Kimi K3 基础上训练的专用模型,号称比 K3 输出 token 减少约 40%,多项外部基准两个付费客户的线上 A/B 测试与自家编码/agent 工作负载上质量持平团队观察到 K3 在生成 token 上有超过 90% 实际用于内部推理,而多轮 agent 工作流会把历史推理完整回灌,导致上下文随轮次近似二次增长Fireworks 跑了 50+ 训练实验和 200+ 评估(在 Serverless Training 平台上完成),保留必要的反思性推理并裁掉冗余循环,跨数学编码指令跟随对话搜索tool use 与软件工程任务做混合训练以让"省 token"行为可泛化在 Specialized Intelligence Index 的 Doxity Bedside Bench 上 Ember-1 在 cost/task 维度对开源/闭源模型(含 GPT-5.6 SolGPT-6 AstraClaude Opus 5)立起新的 Pareto 前沿;Terminal Bench 2.1 82.0%(vs K3 max 80.9%)SWE-bench Verified 92.2%(vs 93.2%)可作为"在不大改模型结构前提下压缩推理开销"的工业样本

English Summary

Fireworks Research shipped Ember-1 on Sep 23, a Kimi K3-specialized model trained to cut ~40% of output tokens while preserving K3 quality across external benchmarks, two customer live A/B tests, and Fireworks' own coding/agent workloads; training used Serverless Training across math, coding, instruction following, conversation, search, tool use, and software engineering tasks. On the Doximity Bedside Bench slice of the Specialized Intelligence Index, Ember-1 sets a new cost/task Pareto frontier over GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5, with Terminal Bench 2.1 at 82.0% (vs K3-max 80.9%) and SWE-bench Verified at 92.2% (vs 93.2%).

为什么值得关注

Fireworks 用 K3 的"省 token"训练让推理成本腰斩,是工业级推理压缩的现成样本

信息源