产品与商业 4.0 · 优秀 2026-08-26 · 文章

Small Models Have Arrived

Segment 联创 Calvin French-Owen 认为小快模型已过实用门槛:gpt-5.6-luna 约 100 tps,跑复杂研究线程仅数十美分,搜几千封邮件也不贵;个性化每日新闻微站评测从 Sonnet 级的约 $1 降到约 $0.10,消费级 AI 定价首次跑得通(此前 $30/月订阅难覆盖推理成本);GLM-5.3 则在帕累托前沿补上新选项他区分两类工作:需前沿模型的 IQ 180 突破型工作(需求会持续复利)与占忙碌创始人约 95% 时间的 token spewer 推进型工作,后者正是快/便宜/够用模型的主场;但企业落地还需新 harness提示注入防护角色与权限体系

打开原文回到归档

Small Models Have Arrived

  • ID: f609f1af
  • 原文链接: https://calv.info/small-models-have-arrived
  • 作者: Calvin French-Owen
  • 日期: 2026-08-26
  • 分类: industry
  • 来源类型: article
  • 标签: small-models, cost, industry, inference
  • 质量评分: 4/5
  • 抓取时间: 2026-08-29T20:21:03+08:00

中文导读

Segment 联创 Calvin French-Owen 认为 gpt-5.6-luna 这类小快模型已跨过实用门槛:他常规看到约 100 tps,模型在他的代码库、邮件和知识库里高速作业;跑相当复杂的研究线程、甚至搜索几千封邮件,API 花费也只有几十美分。他常做的"个性化每日新闻微站"评测(研究他的公开信息、检索 HN/Reddit/Twitter、生成当日头条微站),在 Sonnet 级模型上约需 $1,在 luna 上结果不错且平均成本约 $0.10——消费级 AI 定价第一次跑得通,此前 $30/月订阅很难覆盖推理成本。加上 GLM 5.3,帕累托前沿上又多了一个新选项。他借 Segment 联创 Peter Reinhardt 的观察区分两类工作:"IQ 180" 式突破型工作(前沿模型的需求会持续复利)与 "token spewer" 推进型工作——Peter 同时运营多家公司,自称约 95% 的工作属于后者(开会、催办、推进几十条线),这正是快/便宜/够用模型的主场。企业要真正用起来,还需要新 harness、提示注入防护、角色与权限体系,但他对此有信心。

为什么值得关注

两点值得记录:一是消费级 AI 的成本结构正在翻转,~$0.10/次的个性化生成让订阅制定价首次跑得通,这解释了"为什么还没出现更多消费级 AI 公司"的答案正在失效;二是"忙碌创始人约 95% 时间花在推进型工作上"给出了小模型需求侧的量化直觉——前沿模型与够用模型会长期并存分属两类工作,而 harness、权限与注入防护这类基建是下一个瓶颈。

原文(抓取存档·节选)

# Small Models Have Arrived (AUG 26, 2026, calv.info)

For the past few weeks, I've been playing with gpt-5.6-luna. It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base.

Of course, the biggest thing with luna is the cost. I've tried running some fairly complicated research threads, and it's pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents.

With GLM 5.3, we even have a new option at the Pareto frontier.

One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?" There's a straightforward answer: token costs. [...]

A pet eval of mine is to build a daily news site, personalized to me: research @calvinfo on the internet. figure out what news they might like. build a micro-site with today's top stories, personalized for them. search hn, reddit, twitter, etc.

With the previous generation of models (Sonnet class), you'd spend ~$1 to get anywhere. Charging $30/mo is untenable for a consumer app. [...] But looking at luna, the results are pretty decent, and the average cost is ~$0.10. Now we're talking!

My Segment co-founder Peter and I were recently comparing notes on a hike. Across his various startups, Peter has seen two kinds of work:

1. the "IQ 180" work. some mad scientist genius type comes up with some crazy solution you've never thought of.
2. the "token spewer" work. being ultra responsive, pushing the ball forward across dozens of different fronts.

[...] Peter mentioned that ~95% of the work he does falls into bucket 2. It's hopping on calls. Nudging people. Blocking and tackling.

I think demand for "frontier-level" models is going to keep compounding. Especially for fields that require novel breakthroughs or discovery (engineering, hard science, model training). But I also think the demand for "fast/cheap/good-enough" models is just about to take off. [...]

There's a lot of work that needs to happen to make fast/cheap/good-enough models a reality for business. New harnesses, prompt injection safety, roles, and permissions. But I'm confident we'll figure that out.

(节选省略处用 [...] 标注;完整原文见上方链接。)

Obsidian Notes

  • 内容由 opencli web read --url https://calv.info/small-models-have-arrived --stdout --download-images false -f md 抓取全文生成。
  • 中文导读与价值判断均锚定在条目已有摘要与抓取正文上;未补充正文之外的细节。