Mechanism Design for Alignment and Control
- ID: f894ee7a
- 原文链接: https://arxiv.org/abs/2609.01595
- PDF: https://arxiv.org/pdf/2609.01595v1
- 作者: Dirk Bergemann, Andrew Koh, Stephen Morris
- 日期: 2026-09-01
- 更新: 2026-09-01
- 分类: learning
- 来源类型: paper
- 标签: ai-alignment, mechanism-design, incentive-design, scalable-oversight
- 质量评分: 4/5
- 抓取时间: 2026-09-03T04:29:22Z
中文导读
论文从机制设计角度推广 AI 代理的对齐问题:代理的偏好与能力都不可观测在能力可隐藏不可伪造的单向仿造结构下,获得一个披露原则可行政策的嵌套周期单调性表述以及高阶信念调查多代理的先决条件并应用到提及能力alignment-interpretability 权衡同伴评分纠偏奖励耦合与可扩展监督等具体场景
为什么值得关注
论文从机制设计角度推广 AI 代理的对齐问题:代理的偏好与能力都不可观测在能力可隐藏不可伪造的单向仿造结构下...
关键信息
- 论文标题:Mechanism Design for Alignment and Control
- 作者:Dirk Bergemann, Andrew Koh, Stephen Morris
- arXiv:https://arxiv.org/abs/2609.01595
- 发布时间:2026-09-01
- arXiv 分类:econ.TH, cs.AI, cs.GT
- 关联标签:ai-alignment, mechanism-design, incentive-design, scalable-oversight
English Abstract
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.
English Summary
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure, capabilities can be concealed but not counterfeited, yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。