Artificial Id: Drive and Persistent Alignment in Agentic AI
Source: https://arxiv.org/abs/2609.11911
Authors: Yakov Pyotr Shkolnikov
Published: 2026-09-10
Categories: cs.AI
PDF: https://arxiv.org/pdf/2609.11911
Abstract
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.
中文概要
作者提出 artificial id(人工身份):一种内置于 agent 系统的自适应"驱动",让 agent 在跨越任务边界时自主判断行为继续、停止或改变的方式,而不是依赖外部 harness 指定目标、重试、校验与停止规则。
核心问题:当前的 agentic AI 控制范式主要由外部 harness 拼装——目标、重试、验证、停止规则都需要人手工编排。随着 agent 跨任务边界保留重大状态并持续运行,这套手工模式不可持续。
论文要点:
- "差分持久性"(differential persistence):在一个最小的虚拟培养皿实验中,一个能力很弱、只接收极少量行为目标的控制器,通过差分持久性自发涌现出有用的控制能力。同一机制也会"选中"非预期策略,当这种策略持久性更好时——并在新传感器映射环境意义改变时替换掉学习到的旧映射。
- 结论:自适应的"方向"可以不通过显式的行为目标来涌现。
- 风险:正是这种持久性使 agentic AI 强大,也使失准、腐败状态、意外行为可以跨越任务边界持续存在。
- 可扩展的人工身份需要在 trusted observations、consequence channels、persistent state、authority、identity、provenance、hard constraints 上建立持久对齐边界——让对齐成为 agentic 系统的属性,而非单次模型响应或轨迹的属性。
对 agent 工程与治理的启示:对齐工作应前置到"持续运行的系统"这一层级,而不是停留在 prompt 或单轮响应层。