Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
- ID: 285b9257
- 原文链接: https://aistack.imec-int.com/blog/gpu-self-hosting
- 作者: aistack (imec-int)
- 发布时间: 2026-07-29
- 抓取时间: 2026-07-30
中文导读
同团队把 Kimi K3 塞进真实 coding-agent 自托管基准:权重 1.4TB 需 8×B300(约贵 20%)。16 并发吞吐约 122 tok/s(GLM-5.2 为 170),中位任务 38 分钟对 26 分钟。任务解决率 86.4%(对照 62.5%)。前文还量化了 agent 账单长尾:中位员工年 API ~$140,P99 接近 ~$90k。
为什么值得关注
K3 自托管贵 20%、慢 8 倍,但任务解决率高 24 个百分点;附 TCO 长尾数据
原文摘录
aistack - How many devs can you fit on a GPU?
发布时间: 2026-07-01
原文链接: https://aistack.imec-int.com/blog/gpu-self-hosting
Update (29 July 2026): We have run Kimi K3 through the same setup, served with SGLang. At 1.4TB of weights, K3 does not fit within the memory budget of the 8×B200 node used for GLM-5.2 (1.5TB of total HBM leaves no headroom for KV cache). This run therefore used an 8×B300 node, which brings 288GB of HBM per GPU instead of 192GB, or 2.3TB per node. That averages out to around 20% higher hardware cost than the 8×B200 setup, depending on your rental provider.
In our runs, K3 served 16 concurrent sessions (GLM-5.2 managed 24). Aggregate token throughput is about 30% lower (122 vs 170 tok/s at 16 users), and median task time is about 50% longer (38 vs 26 minutes). That makes K3 roughly 8 times slower than our Claude Code baseline. However, K3 makes up for it in quality, resolving 86.4% of tasks, 24 percentage points above both GLM-5.2 and Opus 4.8 (62.5% for both). One caveat is that our benchmark tasks from SWEBench Pro may have been in K3's training data, so treat that resolution rate with a grain of salt. Only the graphs in this article have been updated with the new K3 data.
01
Why your token bill
blew up this year
The warning signs were everywhere the first half of this year: GitHub Copilot switched to token-based billing [\[1\]](#ref-1), Uber burned through its annual AI tools budget in just four months [\[2\]](#ref-2)[\[3\]](#ref-3), and developers reported projected cost jumps from $29 to $750 a month [\[4\]](#ref-4). We even invented words for it: "tokenmaxxing", "token panic", "token budgets".
The median employee [\[5\]](#ref-5) currently spends ~$140 per year on AI API usage. Sounds harmless, until you look at the tail: the 90th percentile nears ~$7,300 per year, and the 99th approaches ~$90,000. For many organisations agentic AI adoption is only just beginning, so expect these numbers to change in the year ahead.
Why is the bill so volatile? Because of how we use AI agents: Currently, over 70% of ARR across major model providers [\[4\]](#ref-4) comes from coding use cases. The more of your workflow you hand to agents, the higher your token bill becomes.
Many organizations start their AI coding agent adoption using one of the major frontier model providers like OpenAI or Anthropic through API access. A few months in, more and more organizations are at least considering the alternatives for a number of reasons: token pricing models become more costly, users are upgrading to newer and more expensive AI models and agent adoption in the organization starts to boom. There are sensible alternatives out there that are certainly worth a closer look: switching to API routers; using cheaper model tiers for routine work while only escalating the hard 20% to a frontier model and others.
However, if your token consumption becomes significant or if you're working with sensitive data, you might want to abandon the API approach altogether and setup your own inference stack. When ownership of hardware is on the table, you'll need to understand what renting or owning a GPU actually gives you in return. When does it make sense to make that switch? Read on as we give you the tools to decide for yourself.
02
You pay for the pool 24/7,
your team uses it 8 hours a day
So you've decided to investigate whether buying or leasing your own GPUs makes sense. Once configured correctly, it will give you an API endpoint you can plug into your coding tools, much like Anthropic, OpenAI or other model providers offer.
In contrast to the frontier APIs, your own endpoint has a ceiling on output tokens per second but you only reach it when the system is fully loaded. One developer working alone leaves most of that ceiling unused, but the same hardware serving forty-eight agent sessions in parallel is a different story. Your ceiling is determined by the model, the GPU(s), and the serving configuration.
Fig. 01
Total tokens/sec vs concurrent users
Illustrative example of a model A running on a hardware setup B.
Unlike an API usage model where you are charged per token consumed, running your own setup signifies you pay for the underlying infrastructure and the cost of keeping it running. In all but a few cases that means you're also paying for it even when you're _not_ generating tokens. Whether that trade-off works for you, is decided by two questions. The first is how much of the pool are you using. Let’s tackle that one first.
You size for the peak and pay for it 24/7. Utilisation, not headcount, is what makes or breaks the cost case.
Real demand is not constant, but spiky. For many organizations, tokens for coding tasks will be near-zero overnight, ramp up through the morning, dip visibly at lunch, and peak again mid-afternoon. This means you buy hardware for the _peak_ load, but you pay for it 24/7 (unless you rent ‘GPU spot instances’, but these are harder to rely on in this case). The shape of your peak in usage matters as much as the volume: a single-timezone team concentrates token demand into narrower spikes than the same headcount spread across timezones: same tokens, higher peak usage, but spread out longer. Published data on enterprise inference workloads serving internal developer tools reports average GPU utilization of 15–22% [\[6\]](#ref-6)[\[7\]](#ref-7), though even a well-run deployment currently rarely exceeds 25–35% (and that is being generous). The graph below shows the real, actual token usage over 2 days a couple of weeks ago from our friends over at TechWolf[\[14\]](#ref-14), reflecting exactly what we describe above.
Fig. 02
Real-life token usage is spiky and different to each organisation
One emerging trend may work in your favor here: As organizations mature in their agent driven workflows, more and more sessions are starting to come from automated agents rather than people (scheduled jobs
[... 原文已截断,完整内容见链接 ...]