What is the quality of software that AI writes?
- 原文链接: https://www.johndcook.com/blog/2026/08/26/what-is-the-quality-of-software-that-ai-writes/
- 作者: Wayne Joubert
- 日期: 2026-08-26
- 分类: coding
- 来源类型: article
- 标签: ai-coding, code-quality, refactoring, agents, benchmark
- 质量评分: 4/5
原文(抓取存档)
# What is the quality of software that AI writes?
> 作者: Wayne Joubert
> 发布时间: 2026-08-26T18:51:13+00:00
> 原文链接: https://www.johndcook.com/blog/2026/08/26/what-is-the-quality-of-software-that-ai-writes/
---
AI-powered coding agents increase productivity for many developers. But do these agents produce good-quality code?
Some say this doesn’t matter, we are heading toward dark software factories with source code never inspected, and [maybe](https://www.reddit.com/r/singularity/comments/1veslal/elon_musk_the_next_step_is_getting_rid_of_source/) we should [eliminate](https://x.com/elonmusk/status/2084304083851034949) source code altogether—“source code is the new assembly code”.
Others have a different view. Humans sometimes need to debug source code. Source code may need to be audited by humans for compliance. Code should be clear enough for a human to inspect and reason about the algorithms and behavior. Also, good code quality can make the code more legible to agents and reduce unnecessary context.
Developer teams can have many different ideas of what constitutes high-quality software and good coding style. Though there are many valid ways to write code, there is also wide consensus on general [principles](https://blog.codacy.com/code-complexity) of code quality. For example, avoiding very large single functions or modules, avoiding code duplication, avoiding undisciplined feature creep or patchy code and avoiding unnecessarily deep class hierarchies or function call chains. Some coding style choices are testable empirically for impact on developer productivity. Furthermore, some code complexity measures can be computed objectively and programmatically.
My experiences are with GPT 5.5 (Extra High reasoning) and 5.6 (Extra High, Max and occasionally Ultra). Much of my experience is “out of the box” usage of Codex, with simple AGENTS.md file, though I am working on [improving](https://openai.com/index/harness-engineering) the engineering of guidance files, and it is helping. My source code is mostly Python. Unfortunately it is difficult to generalize any one set of experiences universally, since developers have different code bases, languages, models, harnesses and AGENTS.md files. One-shotting a simple computer game or website would be very different from developing a complex research code in a new domain.
At first glance, the AI-written code is not incomprehensible. It does not use odd variable names like “iiii” or “a87275,” and it does not look like it came out of an [obfuscated code competition](https://www.ioccc.org/). But, in my experience, still the generated code has deficiencies:
- The agent has a tendency to write much more code than is necessary (commonly 2-3X more—see also related findings [here](https://arxiv.org/abs/2603.24755)). Though it is capable of deleting code, its primary impulse seems to be to write more code.
- You can work with the agent to shorten the code, but it takes work. The agent it seems is not fluent in finding structural simplifications and then extracting commonalities. In one session I spent 1/2 hour having the agent write a few hundred lines of code, and 4 hours to get it to shorten and simplify the code. You can imagine the kind of technical debt this would accumulate.
- It behaves as though code simplification is much more out of its reach than code generation. At times it just completely fails to do some simplification task I ask it to do.
- It has no instinct for when to break a file into multiple files for conceptual clarity, even if a file becomes over 10,000 lines long.
- It can reinvent a similar but different helper function in different code modules rather than designing a simple reusable function once.
- It can make massive function argument lists with 10-20 arguments rather than recognizing that the parameters may form a coherent concept representable as an abstraction or parameter object.
- Importantly, it can define functions based on abstractions that do not model the underlying domain well and are hard to decipher. When I called it on this, it said: “You’re right. The code is naming implementation mechanics instead of stating intent … it forces the reader to mentally execute several layers of infrastructure just to discover that it means.”
- It often invents terminology that cannot instantly be understood by the reader (the source code analogy of [Don’t Make Me Think](https://sensible.com/dont-make-me-think/)).
- It can repeat the same expression multiple times instead of defining a variable with a meaningful name to represent the quantity.
- It can hardwire unexplained “magic constants” into the code instead of defining them with meaningful names.
Indeed, when pressed, the models are sometimes capable of doing better. For example, for a hard design problem, 5.6 Ultra was capable of creating a good object design that was a good match to the problem domain, when I asked it to look hard at the problem—better than the less sophisticated models.
I would certainly expect that with more engineering of the agent guidance files, many or most of these problems would get better. However, it should not require extreme measures to get coding agents to write good code.
I have not compared other coding agents, but it would not surprise me if they had similar issues. Rightly, the coding models have been optimized for their software development utility, and this has undoubtedly succeeded in a revolutionary way.
It seems there is not yet a widely accepted code-quality benchmark playing the role for frontier coding models that SWE-bench has played for software engineering capability (though there are [efforts](https://labs.scale.com/leaderboard/sweatlas-refactoring)). It’s especially interesting because many aspects of code quality are verifiable, making the problem seemingly quite amenable to treatment in post-training. I am hoping that someone can put together a good benchmark for this problem, and that the frontier labs will embrace these kinds of evaluations in model development.
中文导读
作者用 GPT 5.5/5.6 在 Codex 上写 Python 的亲身经历列出 AI 生成代码的 10 条结构性问题,结论是「生成远比简化擅长」:agent 写出的代码量通常是必要值的 2-3 倍;半小时写几百行、再花四小时让它简化很常见;不会主动拆分超 10,000 行的单文件;不同模块重复造相似但不同的 helper;动辄 10-20 个参数的函数签名;抽象常与底层领域不匹配——被点破时模型自认「在命名实现机制而不是表达意图」。一个有效动作:让 GPT 5.6 Ultra 盯一个困难设计问题,能给出匹配领域的好对象设计,说明压它做更深思考有效。判断:现在缺一个像 SWE-bench 那样被广泛接受的代码质量基准,而代码质量的很多维度可以程序化衡量,正是 post-training 的好靶子。
为什么值得关注
AI 代码质量的十条实证:生成是简化的 2-3 倍、万行单文件不拆、参数列表膨胀;缺的是 SWE-bench 级的代码质量基准
English Summary
Based on hands-on experience with GPT 5.5/5.6 writing Python on Codex, the author lists ten structural problems of AI-generated code, concluding generation is far better than simplification: agents write 2-3x more code than needed; half an hour to generate hundreds of lines followed by four hours of simplification is common; they won't split 10,000+ line files unprompted; modules duplicate similar-but-different helpers; function signatures balloon to 10-20 parameters; abstractions mismatch the domain -- and the model admits it is 'naming implementation mechanics rather than expressing intent' when pressed. Pushing GPT 5.6 Ultra on a hard design problem does yield domain-appropriate object designs. The gap: no widely accepted SWE-bench-equivalent for code quality, though many quality dimensions are programmatically measurable -- a good post-training target.
Obsidian Notes
- Body fetched via
opencli web read --url https://www.johndcook.com/blog/2026/08/26/what-is-the-quality-of-software-that-ai-writes/ --download-images false -f md(2026-08-27). - 原文全文存档于上方代码块;导读锚定正文事实。