Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents
Source: https://arxiv.org/abs/2608.11110
Author: Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram
Original date: 2026-08-11
Added by: AAIF daily-intake-evening 2026-08-13
摘要
论文把 tool-using agent 的跨语言能力从最终回答转移到 action policy 本身来测量:8 个模型、6 个 benchmark、41 种语言、2.38M rollouts。结果显示 frontier 模型在 greedy decoding 下也只保留约 71-73% action policy,且英文 pivot 对非英文任务是 causally load-bearing。
English Summary
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.
入库理由
- quality_score: 4
- category: agents
- tags: tool-using-agents, cross-lingual, policy-retention, evaluation
- one_liner: 跨语言 agent 评测应看 action policy 是否保留,而不只看最终自然语言回答。
Obsidian evidence excerpt
t,单纯 query-level 路由不足以做决策。
来源:https://arxiv.org/abs/2608.09155
17. **Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents** — arXiv:2608.08389。Deep research agent 的边际价值递减问题:context 快速增长,额外 evidence 的边际价值下降。给出 marginal value estimation 让 agent 学会"在哪个点停止 retrieval"。
来源:https://arxiv.org/abs/2608.08389
18. **Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents** — arXiv:2608.11110。把 tool-using agent 的 action policy 当成 measured object(而不是 final answer),8 模型 / 6 benchmark / 41 语言 / 2.38M rollouts。Greedy decoding 下四个 frontier 模型仍只能保留 71-73% action policy 跨语言;模型身份只解释 5.7% variance;<10B 参数下完全失效;agents 用英文 pivot 处理非英文任务——pivot 是 causally load-bearing,模型不放弃。
来源:https://arxiv.org/abs/2608.11110
19. **ASCon: A Direction-Aware Reciprocal Agent–Step Contextualization Model for Failure Attribution in Multi-Agent Systems** — arXiv:2608.10646。MAS failure attribution 回答"who/when/why"——定向 + 双向 agent–step 上下文建模。
来源:https://arxiv.org/abs/2608.10646
20. **Automating and Scaling Behavioral Scientific Research on AI Agents (AEROBAT)** — arXiv:2608.10030。第一个 multi-agent system 用自动化方式做 AI agent 的 behavioral scientific research。
来源:https://arxiv.org/abs/2608.10030
## 值得精读的论文
1. **SkillZip(2608.11079)**
精读价值:今天"memory 范式切换"主线里工程化程度最高的一篇。self-evolving agent 长期痛点是 skill 越长越贵、越难维护,但 generic prompt compression 不适合 skill 结构(contract / workflow / tool / exception 各有语义)。SkillZip 的 typed MDL 目标 + Zip-on-Write 模式给出了"边演化边压缩"的具体实现,而且 evaluation-free(不靠 rollout 验证),在 Hermes 自身的 skill / memory 演进机制里直接可借鉴——Hermes 的 profile skill 超过一定阈值后也需要类似的结构化压缩。需要细看的是 typed MDL objective 的 sharing threshold 是怎么 derived、coverage constraint 在 partial extraction 下是否还成立、以及 contract-type 集合的完整度。
2. **MAP-Graph(2608.10509)**
精读价值:把 provenance 从"事后审计"升级为"运行时 control signal"。permission filtering + path trust + risk-sensitive gate 三层组合在 2,700 合成任务上 94.96% success,对 Not an A11y 这类 prompt injection 攻击场景是直接相关的工程参考——如果 agent 的每一步 action 都要穿过 provenance 校验,攻击的成功面就大幅收窄。需要细看的是 path trust 的 multiplicative 形式在长 lineage 下的数值稳定性,以及 risk-sensitive gate 的 action risk 评分实操。
3. **ReRo
arXiv metadata / abstract
- arXiv id: 2608.11110
- authors: Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram
- published: 2026-08-11
- updated: 2026-08-11
- categories: cs.CL
- PDF: https://arxiv.org/pdf/2608.11110v1
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.