产品与商业 4.0 · 优秀 2026-07-18 · 文章

AI Mania Is Eviscerating Global Decision-Making

文章不是泛泛批评 AI,而是拆解企业 AI 狂热如何改坏决策系统:内部 chatbot 没人用客户 chatbot 失效却不算事故员工为了 token 指标伪装成高强度使用者作者把问题落到组织政治:当客户高层也公开宣称 100x 生产力时,供应商和内部管理者很难讲真话,AI 采用率容易变成表演指标

打开原文回到归档

AI Mania Is Eviscerating Global Decision-Making

摘要(中文)

文章不是泛泛批评 AI,而是拆解企业 AI 狂热如何改坏决策系统:内部 chatbot 没人用客户 chatbot 失效却不算事故员工为了 token 指标伪装成高强度使用者作者把问题落到组织政治:当客户高层也公开宣称 100x 生产力时,供应商和内部管理者很难讲真话,AI 采用率容易变成表演指标

Summary (English)

Nik Suresh argues from consulting and enterprise experience that AI mania is distorting incentives, metrics and truth-telling inside large organizations.

One-liner

企业 AI 狂热最危险的地方,是奖励看起来用了 AI而不是真实交付

原文 / 元数据抓取

AI Mania Is Eviscerating Global Decision-Making

发布时间: 2026-07-18
原文链接: https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/

AI Mania Is Eviscerating Global Decision-Making

Published on July 18, 2026

Note: This has been cross-posted to my company's blog, in case you think there is some use in sharing with someone in a format that looks more authoritative. Link here.

I strongly believe there are entire companies right now under heavy AI psychosis and it’s impossible to have rational conversations with them about it. I can’t name any specific people because they include personal friends I deeply respect, but I worry about how this plays out.
Mitchell Hashimoto, of HashiCorp and Ghostty fame

Over the past year, I’ve run point on all of our company’s sales, led the technical components of all but two of our engagements, and over the lifetime of this blog have had something like 300 catchups with professionals from around the world. This has ranged from people on the ground in niche service industries to executives at Fortune 500 companies1. Because of this, I've had a front-row view to our collective institutions across both the private and public sector undergoing breath-taking mass psychosis. This essay is an attempt to describe the bizarre dynamics that are currently at play, as I am in the rare position where my wellbeing is not contingent on paying lip service to madness, and to reassure the people trying to survive amidst all of this that they are not crazy.

The reality is thus: the people in charge either have no plan, or see no path forwards other than keeping their heads down. Not at banks, not at hospitals, not in our government institutions. The world’s organisations have been captured by people in the throes of frothing excitement, and saner people who now live in a state of constant commingled fear and frustration.

I. AI Investments Are Generally Total Failures

Reading this while working for a division that pivoted to provide interfaces for agentic workflows, only to discover that only ten users had ever touched the products we made for agents, only to pivot again to support for agentic workflows, which has a lot of competition because every company has to do something agentic now and there's only like four things you can do in that space, is bracing.
– An editor of this essay

Are companies actually seeing massive productivity gains from their AI adoption? Does any of this sordid affair _make sense_?

This should be an easy question, but it is surprisingly hard to get a straight answer to it. Executives that tell the press that their company has gone insane will quickly find themselves removed from their positions. Employees who are honest will find themselves fired in short-order, or “randomly” selected for a round of layoffs. In fact, it is in the interests of almost every actor in the space – boards, executives, employees, vendors, consultants – to obfuscate and misrepresent the success rate of AI projects. Many publicly traded companies are putting out announcements about their AI productivity gains when I know for a fact that the businesses have done nothing other than purchase Copilot licenses and declare victory.

Yet we need to know if these projects are panning out – if the total focus on AI as a core tenet of business strategy is succeeding at a reasonable rate, then a discussion about the relative risk and reward is warranted.

Unfortunately, we live in a dark timeline. All of the AI projects we have observed as a team are failing. Every single one – we have seen 0% success in a year and a half, not only amongst projects we have been asked to participate in2, but even within projects that we have observed in passing while doing totally unrelated work. Even if you grant that AI tooling accelerates specific workloads, the method and scale of the current investments is senseless. Frequently the failure is not related to AI itself, but rather that companies are terminally bad at running software projects effectively, and as I have remarked previously, AI projects are subject to all the failure modes of normal projects _plus_ you can get everything right and then still fail because of the method's novelty. Very few companies are so good at shipping software that they can afford the extra risk profile.

Often enough, though, it’s an actual failure in what LLMs can accomplish. The most common version of this, being rolled out across businesses around the world, is the internally-facing chatbot, or for the more daring company, the customer-facing chatbot. The story is always the same. For the former, I’ve never seen substantial internal uptake from inside a business. Employees don’t use internal chatbots because companies tend to have low-quality documentation and an LLM is not psychic – it can only know things that have been written down and made accessible. For the latter customer-facing applications, I have rarely had a pleasant experience as a consumer, with _perhaps_ the exception of live transcription during medical appointments – hardly something worth pivoting an entire organisation around. In both cases, project leaders are very careful to avoid tracking basic metrics, such as whether the tools are being used at all, or they track metrics that are easily gamed.

For example, my last consumer interaction was attempting to get help from Mitsubishi following an automotive failure, where a very polite robot asked me to describe the problem and that I’d receive a call back as soon as someone was available. This was the single most competent implementation of such a project I’ve seen in the wild, in that the voice was natural sounding, responded quickly, was clearly “live” in production, and promised a swift resolution.

That was six months ago, and I did not, in fact, get a call back.

When Mitsubishi did not call me back, what happened? Did that request just go into the void, showing one less incident for the year? Does it appear that the phone bot resolved my query without the need for human intervention? All we know is that it didn’t show up as an error, or I’d have received a call. I’m sure it looks great in all sorts of ways except the one that matters, which is that I was planning to buy a car and decided _not_ to buy another one of theirs.

For this reason, our team has quickly learned while on an engagement not to ask anything about ongoing AI projects in any context – by the time that project has started, it is too late for the management team, and intervention is not possible until a crisis point is inevitably reached. There is no conceivable positive outcome. The failure rate is so high that even basic inquiry leaves us in an untenable position. Any coherent question about how it’s going, what the goal is, who is using it, constitutes an inadvertent attack on the chain of command responsible for the work because _there are no good answers to anything_. Even in rare cases where my interlocutor has stated that things are going well (usually while the project is still mid-flight and failure has not had a chance to manifest), it is generally obvious that they are doomed, but at least in these cases I can simply agree and then go home to scream into a pillow for six hours straight3.

All of this is to say that I am very confident that almost every report at a company about “massive AI productivity gains” is untrue as a matter of brute fact. Even if some companies are seeing clear gains, this is the exception, not the norm. With that assumption in place, we can talk about the dynamics at play, and how it has become impossible for many organisations to stay focused on things that actually matter to their long-term (or even short-term) health.

II. Heretics Will Be Shot

It has become outright dangerous to even raise the possibility that AI might not be the solution to a problem, let alone be the sole focus of a company’s entire strategy.

In every sufficiently large business we have observed (say, with 500+ employees), we have noted that continued advancement, and increasingly continued employment, has started to require repeated professions of belief in the transformative power of AI for said business. I am not talking about providing ideas about how to use AI in the business – I mean _religious_ profession, declarations of faith. Overwhelmingly these statements are made by non-technicians, though it is not uncommon for technicians to emit deranged statements to curry favour.

There have been several occasions where I have seen someone, apropos of nothing, blurt out almost word-for-word “AI is changing everything”, only to concede moments later that their organisation does not currently use LLMs for anything, and indeed, that they cannot name a single thing that has changed other than they get some use out of ChatGPT (frequently the free-tier). In one extreme case, I have seen an executive confess that they had never even used ChatGPT or any AI tool in their life, immediately after producing a technical strategy for an organisation with $2B+ in revenue which was entirely centered around AI.

Initially these statements were so absurd on their face that I thought it was some cynical ploy to achieve thought leader status, and there are certainly some people doing this – I have had it admitted to me. But the broader reality is so much worse: people who have no background in the technology at all _actually believe what they are saying_. As a general rule you should avoid getting into business with a liar, but if you _must_, you can at least reason with them even if only in private. A true believer is much more threatening because they are impervious to even inducement by self-interest.

The turning point in my belief was watching someone with a spectacular amount of money on the line fire their highest performers because they were achieving that performance without LLMs. When an employer _publicly_ talks about AI innovation, we have to ask ourselves if they’re simply trying to manipulate the market or customers. When they _privately_ commit to strategies like this with their own money at stake, with no attempt to communicate that strategy to external clients, I can only assume they really mean what they’re saying.

A while ago, I wrote “Contra Ptacek’s Terrible Article On AI”, which was focused on the fact that many of Ptacek’s points in his own essay “My AI Skeptic Friends Are All Nuts” were internally inconsistent4. But on the crux of the matter, we are actually in total agreement, because he opens his essay with this:

Tech execs are mandating LLM adoption. That’s bad strategy.

Which is to say that we can sidestep arguments about the precise utility of LLMs entirely and we’re left in a very simple place – it is entirely obvious to both myself and Ptacek, two people that are coming at this from fairly opposed views, that people are being really, really stupid about this, _and_ that organisations are demanding bizarre workflow constraints from their specialist staff.5

These mandates have led to extremely strange places. Several of my peers now “AI-wash” their work, meaning that even when they can perfectly competently execute on their jobs to the satisfaction of their management teams, said managers are unhappy if the engineers haven’t used AI in the work… so now they’re lying about using LLMs even in contexts where their professional judgement is that they aren’t the appropriate tool. They just do the work, the same way they have for decades, and say Claude did it. Others are being measured on their AI

Obsidian intake evidence excerpt

AK-RSS-Digest(89源精选) · 2026-07-20

  • status: completed
  • Obsidian: /Users/gracker/Library/Mobile Documents/iCloud~md~obsidian/Documents/Obsidian/OpenClaw定时任务/AK-RSS-Digest(89源精选)/2026-07-20-AK-RSS-Digest.md
  • Evidence: /Users/gracker/.hermes/evidence/ak-rss-digest/2026-07-20
  • 覆盖范围: 89 个 RSS 源,最近 7 天候选 50 条;3 个 feed 抓取失败,未影响本期精选。

今日精选

1. AI Mania Is Eviscerating Global Decision-Making(8.9/10) 2. Overtraining as the path to human-like AI(8.8/10) 3. What’s the deal with all the random weekly quota resets for agents lately?(8.4/10) 4. Art Doesn’t Scale(8.3/10) 5. Plumbing Homebrew into the vulnerability ecosystem(8.2/10)

可发布正文如下:

评分:8.9/10 推荐语:这篇不是反 AI,而是在拆企业 AI 狂热怎么把决策系统改坏。最有价值的是一线咨询视角:指标被游戏化、项目不敢问效果、供应商和客户高层互相绑架,很多公司最后买到的是“贴了 AI 标签的普通项目”。 摘要:作者基于咨询和企业沟通经历,写出大组织在 AI 投资上同时缺少真实指标和纠错机制的状态:内部 chatbot 没人用、客户 chatbot 失效却不算事故、员工为了 token 指标伪装成高强度使用者。文章后半把问题落到组织政治:当客户高层也在公开宣称 100x 生产力时,供应商和内部管理者都很难讲真话。 链接:https://ludic.mataroa.blog/blog/ai-mania-is-eviscerating-global-decision-making/

  • 标题:AI Mania Is Eviscerating Global Decision-Making

评分:8.8/10 推荐语:这是对 Gwern “catapulting/grokking”思路的一次清楚转述,适合跟踪下一轮大模型训练路线的人读。它把争议点说得很窄:如果当前模型缺的是更深层泛化,问题可能不在继续堆数据,而在让超大模型在较小数据集上被迫压出更简洁的规则。 摘要:文章先解释 grokking:模型在训练损失归零后继续训练,可能从记忆样本跳到掌握底层规则。作者随后讨论反向训练策略的代价:用百兆亿参数级模型、小数据集和长时间训练去赌一次能力跃迁,技术风险和组织耐心都是门槛。 链接:https://seangoedecke.com/overtraining-as-the-path-to-human-like-ai/

  • 标题:Overtraining as the path to human-like AI

评分:8.4/10 推荐语:这篇把 coding agent 订阅制的一个小现象写成了产品策略观察:频繁重置额度表面上是福利,实际会改变重度用户的使用节奏和付费判断。它还指出一个竞争信号:在 Fable 5、GPT-5.6 Sol、Grok、Muse、Kimi 同时挤压市场时,额度重置也可能是防止用户试用竞品的留存手段。 摘要:作者记录了 OpenAI 在两周内多次重置 Codex 周额度,以及 Anthropic/OpenAI 在新模型发布期对额度策略的调整。文章的价值不在抱怨免费额度,而在说明 agent 产品的计费、算力调度、用户心理和竞品试用窗口已经绑在一起。 链接:https://minimaxir.com/2026/07/agent-quota-reset/

  • 标题:What’s the deal with all the random weekly quota resets for agents lately?

评分:8.3/10 推荐语:这篇适合放在 AI 生成内容争论里反复引用:作者不争“AI 作品能不能让人有感觉”,而是把焦点放到创作上下文、署名、等待和稀缺性。它的判断很直接:很多人珍惜作品,不只因为结果好看,也因为有人花了时间和代价把它做出来。 摘要:作者从漫画、赝品和 AI 续作谈起,说明作品价值包含创作者经历、完成过程和读者对真实来源的信任。文章后半把“效率”问题讲清:AI 可以填满空白,但如果平台默认用廉价生成物替代人工创作,用户失去的是选择权。 链接:https://matduggan.com/art-doesnt-scale/

  • 标题:Art Doesn’t Scale

评分:8.2/10 推荐语:这是一篇高质量工程复盘,讲清 brew vulns 从个人 gem 进入 Homebrew 6.0.11 的完整路径。它不只是功能发布,而是把包标识、OSV 生态注册、补丁声明、版本比较、advisory database 和 CI 边界全部摊开。 摘要:作者先说明 Homebrew 扫描 CVE 的难点:上游 repo 和版本号不够,很多 formula 带补丁,OSV/NVD 数据也经常缺少可查询的 package 信息。后文给出落地路径:新增 patch resolves 元数据、注册 pkg:brew 和 OSV Homebrew ecosystem、生成 Homebrew advisory database,并把扫描器合入 brew 主仓。 链接:https://nesbitt.io/2026/07/17/plumbing-homebrew-into-the-vulnerability-ecosystem.html

  • 标题:Plumbing Homebrew into the vulnerability ecosystem

可直接发布文案

本期 AK RSS 里有几篇信号很强:一篇拆企业 AI 狂热怎么把决策机制搞坏,一篇转述 Gwern 的 grokking/overtraining 路线,一篇从 agent 额度重置看订阅产品策略,还有两篇分别谈 AI 生成内容的“来源价值”和 Homebrew 漏洞扫描的工程落地。

最推荐先读《AI Mania Is Eviscerating Global Decision-Making》:它不是泛泛吐槽 AI,而是写清了大公司里为什么越来越难诚实讨论 AI 项目的效果。接着读《Overtraining as the path to human-like AI》,换到模型训练路线层面,看下一次能力跃迁可能押在哪。

AI #Agent #LLM #软件工程 #RSS精选

备选短文案

  • 企业 AI 狂热最危险的地方,不是用了多少 LLM,而是它开始奖励错误指标、惩罚诚实反馈。今天 RSS 里这篇 Ludic 文章很值得读。
  • Gwern 的 catapulting/grokking 思路被 Sean Goedecke 转述得很清楚:如果模型缺的是更深层泛化,继续堆数据可能不是唯一解。
  • brew vulns 这篇复盘很好看:一个小工具进入 Homebrew 主线,背后要补齐 purl、OSV、patch metadata、版本比较和 advisory database。

本期未入选说明

  • Claude Code 使用 Rust 版 Bun:证据有趣,但正文偏短,适合作为技术笔记,不够本期精选强度。
  • Make It Work vs. Make It Good:观点顺手,但更像短随笔,未过 8 分线。
  • OpenAI 产品重组:信息有用,但原文发布时间较早,且更接近新闻报道,本期不放入精选。