Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
Source: https://withspecific.com/benchmarks/real-swe
Author: Specific
Published: September 2026
Overview
Real-SWE evaluates frontier coding agents on tasks lifted from licensed, private production codebases of real companies — code and solutions that don't exist on the public internet. Tasks have real business consequences (getting billing right, calculating taxes, migrating customers) and must respect company-specific conventions. The question: can a coding agent actually do the work of a software engineer in the real world?
Design
- Natively out-of-distribution: "99% of tokens in real-world enterprises are hidden away from the frontier models." Tasks are actual engineer-assigned work, not expert-authored puzzle tasks.
- Native harnesses: model-and-harness combinations (Claude Code, Codex CLI, Gemini CLI...), reflecting how enterprise engineers actually work. Resolution rate = pass@1 averaged over 8 independent runs per task with 95% CIs.
- Harder task shape than peers: median instruction 1,742 chars (FrontierCode 2,056, DeepSWE 1,975, Terminal-Bench 3 1,584); reference solutions edit a median of 11 files vs 6 for FrontierCode and DeepSWE. Sample codebases include a Luma/Partiful competitor with 200K+ users and a consumer fintech platform processing 100K+ bank statements.
Leaderboard (September 2026)
| # | Model | Harness | Resolution | |---|-------|---------|-----------| | 1 | Claude Fable 5.1 | Claude Code | 38.8% | | 2 | GPT-6 Astra | Codex CLI | 33.8% | | 3 | Gemini 3.8 Flash | Gemini CLI | 31.2% | | 4 | GLM 5.3 | Claude Code | 28.8% | | =5 | Grok 4.6 | Grok Build | 23.8% | | =5 | Muse Spark 1.3 | Muse Code | 23.8% | | 7 | Kimi K3 | Kimi Code | 18.8% | | 8 | GPT-5.6 Sol | Codex CLI | 16.2% |
Findings
- 6 of 10 sample tasks have resolution rates below 15%. Hardest: Tax jurisdiction 3.1%, Linearizable scan 4.7%, S3 datastore measurement 10.9%. Easiest: Multi-region sweep 67.2%, API keys & environments 65.6%.
- Failure is not about long horizons: 71.4% of rollouts under 10 minutes failed vs 73.4% of longer ones. Triaging multiple systems and understanding requirements in codebases riddled with existing business logic is the difficulty — not task duration.
- Models miss requirements and don't verify assumptions, and are weaker at company-specific coding patterns: "Many enterprises care about code standards and patterns... we're far from that reality."
中文概要
Specific 发布 Real-SWE 基准:任务来自经授权的真实企业私有生产代码库,代码与解法不在公共互联网上,智能体必须处理计费、税务、客户迁移等有真实业务后果的多服务变更,并遵守每家公司自己的编码规范。首期榜单(每任务 8 次独立运行取 pass@1 均值):Claude Fable 5.1 以 38.8% 居首,GPT-6 Astra 33.8%、Gemini 3.8 Flash 31.2%、GLM 5.3(Claude Code)28.8%。任务形态显著更难:参考解法中位改动 11 个文件(FrontierCode/DeepSWE 为 6),提示词中位 1,742 字符。一个反直觉发现:10 分钟内的短任务失败率 71.4%,与长任务(73.4%)几乎相同——瓶颈在理解既有业务逻辑与跨系统排查,而非任务时长。最难的"税率管辖"任务通过率仅 3.1%。结论:私有代码天然分布外,企业级编码规范遵循与需求验证仍是前沿模型的集体短板。