模型与实验室 4.0 · 优秀 2026-06-07 · X

Anthropic 开源对齐工具 Petri 捐赠给 Meridian Labs:版本 3.0 更新

2025 年 10 月,我们发布了 Petri,这是一个可用于任何大型语言模型的开源对齐测试工具箱Petri 诞生于 Anthropic Fellows 计划,可用于快速便捷地测试 AI 模型在欺骗谄媚和对有害请求配合等令人担忧的倾向上它是我们开发开放且对整个 AI 社区有用的对齐工具的努力的一部分 自 Claude Sonnet 4.5 以来,Petri 一直是每个 Claude 模型对齐评估的一部分它通过一个独立的"审计员"模型模拟一系列对齐相关场景,比较新模型的行为表现然后一个"裁判"模型对产生的对话记录进行评分,识别对齐偏差行为 我们很高兴看到外部组织也在使用 Petri:例如...

打开原文回到归档

Anthropic 开源对齐工具 Petri 捐赠给 Meridian Labs:版本 3.0 更新

English (Original)

Alignment

Donating our open-source alignment tool

May 7, 2026

In October 2025, we launched Petri, an open-source toolbox of alignment tests that can be applied to any large language model. Petri, which was developed as part of our Anthropic Fellows program, can be used to rapidly and easily test AI models for concerning tendencies like deception, sycophancy, and cooperation with harmful requests. It’s part of our efforts to develop alignment tools that are open and useful for the whole AI development community.

Petri has been part of our alignment assessment for every Claude model since Claude Sonnet 4.5. It compares how the new model behaves across a range of alignment-relevant scenarios that are simulated by a separate “auditor” model. A further “judge” model then scores the resulting transcripts for misaligned behaviors.

We’ve been pleased to see Petri being used by external organizations: for example, the UK’s AI Security Institute (AISI) made it a major part of how they evaluate models for their propensity to sabotage AI research.

We’re now updating Petri to its third version. Here are some of the biggest changes:

  • _Adaptability._ Petri 3.0 involves major architectural changes that allow users to adapt it to more uses, in particular by splitting the auditor model and the target model into separate components that can be tweaked separately;
  • _Realism._ Despite the fact that alignment researchers try to make tests appear realistic, a model can often deduce from various artificialities in the setup that it’s actually part of a test. And if the model is aware it’s being evaluated, the researcher is no longer able to see how the model behaves _in general_. An add-on to Petri, which we’re calling “Dish,” makes the setup far more realistic, for example by running the tests using the model’s real system prompt and the real “scaffold” (the software that wraps around the model to help it meet its goals) that would be used in genuine model deployments;
  • _Depth_. We’ve now integrated Petri with our other open-source alignment tool, Bloom, which can perform much more in-depth assessments of specific chosen behaviors (in comparison to Petri’s wider-ranging approach).

We’re also giving Petri a new home. We have handed over its development to Meridian Labs, an AI evaluation nonprofit. This move—similar to when we donated the Model Context Protocol (MCP) to the Linux Foundation—will help ensure that Petri remains independent of any AI lab, so that its results will be seen as neutral and credible by those across the industry and beyond.

As part of Meridian Labs, Petri joins other tools like Inspect and Scout, building a technology stack that is open to labs, independent researchers, and governments alike, at a time when reliable tests of AI model behavior matter more than ever.

You can read more about Petri 3.0 on the Meridian Labs blog.

Instructions to install and use Petri can be found on the Petri website.

https://twitter.com/intent/tweet?text=https://www.anthropic.com/research/donating-open-source-petrihttps://www.linkedin.com/shareArticle?mini=true&url=https://www.anthropic.com/research/donating-open-source-petri

Related content

Paving the way for agents in biology

Read more

Making Claude a chemist

Read more

Coding agents in the social sciences

Results from a survey of 1,260 social scientists about AI and coding agent use.

Read more

中文摘要

Anthropic 将开源对齐测试工具 Petri 3.0 捐赠给 Meridian Labs(一家 AI 评估非营利组织),以确保其独立于任何 AI 实验室。Petri 3.0 在架构上把审计员模型与目标模型拆开,新增 Dish 组件让测试设置更贴近真实部署,并与 Anthropic 另一款开源对齐工具 Bloom 集成做更深入的特定行为评估。

中文翻译

2026 年 5 月 7 日,Anthropic 宣布将开源对齐工具 Petri 捐赠给 Meridian Labs。

2025 年 10 月,Anthropic 发布了 Petri——一个可应用于任何大型语言模型的开源对齐测试工具箱。Petri 诞生于 Anthropic Fellows 项目,可用于快速便捷地测试模型在欺骗、阿谀奉承、与有害请求配合等令人担忧倾向上的行为。这是 Anthropic 开发开放、有益于整个 AI 社区的对齐工具努力的一部分。

自 Claude Sonnet 4.5 起,Petri 已成为每款 Claude 模型对齐评估的一部分。它通过独立的"审计员"模型模拟一系列与对齐相关的场景,比较新模型的行为表现;再由"裁判"模型对生成的对话记录按对齐偏差行为评分。已有外部机构使用 Petri——例如英国 AI Security Institute(AISI)将其作为评估模型"蓄意破坏 AI 研究倾向"的主要工具。

Petri 3.0 更新要点

适配性:架构上做了重大重构,把"审计员模型"与"目标模型"拆成两个可独立调整的组件,便于拓展更多用途。

真实感:尽管对齐研究者尽力让测试显得真实,但模型往往能通过实验设置中的人造痕迹识别出自己在被测。一旦意识到在被评估,模型就不再展现它一般情况下行为。新增的 Dish 组件让测试设置更贴近真实部署场景——使用模型真实的 system prompt 与真实的 scaffold(包裹模型的运行软件)。

深度:Petri 现已与 Anthropic 另一款开源对齐工具 Bloom 集成。Bloom 可对特定选定行为做更深入的评估,弥补 Petri 覆盖广但深度有限的不足。

新的归属——Meridian Labs

Petri 的开发权正式移交给 Meridian Labs——一家 AI 评估非营利组织。这一举措与此前 Anthropic 将 Model Context Protocol(MCP)捐赠给 Linux Foundation 的做法相似,有助于确保 Petri 独立于任何 AI 实验室,使其结果在整个行业及更广范围内被视为中立可信。

加入 Meridian Labs 后,Petri 与 Inspect、Scout 等工具共同构建一套对实验室、独立研究者、各国政府开放的技术栈。在 AI 模型行为可靠测试比以往任何时候都更重要的当下,这套技术栈尤为关键。

Petri 3.0 的更多细节可阅读 Meridian Labs 博客文章。安装与使用说明见 Petri 官网。

*本文件由 AAIF Content Fetcher 自动抓取并双语整理。原文版权归原作者所有。*