模型与实验室 3.0 · 值得看 2026-06-23 · 文章

ChatGPT Images 2.0 发布,Where's Waldo 风格测试引发争议

ChatGPT Images 2.0 发布,Where's Waldo 风格测试引发争议

打开原文回到归档

ChatGPT Images 2.0 发布,Where's Waldo 风格测试引发争议

Source: Simon Willison | 2026-04-21 URL: https://simonwillison.net/2026/Apr/21/gpt-image-2/

Summary (Chinese)

OpenAI 发布 ChatGPT Images 2.0,Sam Altman 称从 gpt-image-1 到 2 是巨大飞跃。Simon Willison 测试发现:细节还原很好但文字渲染仍有错误;让模型找自己生成的 raccoon 并画红圈,模型答错了自己在图里画的内容——说明多模态模型的自我验证能力仍存在明显漏洞。这类 Where's Waldo 风格测试暴露了当前图像生成+视觉推理 pipeline 的薄弱环节:模型对自己生成内容的视觉验证能力有限,在需要高精度视觉推理的场景需要谨慎使用。

Summary (English)

OpenAI releases ChatGPT Images 2.0. Simon Willison's testing reveals good detail reproduction but persistent text rendering errors. More concerning: asking the model to find a raccoon it drew and circle it in red resulted in wrong answers — exposing that multimodal models still have significant self-verification weaknesses.

注:原文抓取失败(web_fetch 被封锁),此 content 文件基于 AK RSS Digest 摘要整理。原始内容请参考 AK RSS Digest 源文。