CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
- ID: fc2c84f3
- 原文链接: https://arxiv.org/abs/2609.01600
- PDF: https://arxiv.org/pdf/2609.01600v1
- 作者: Damien Sileo, Dimitri Kachler
- 日期: 2026-09-01
- 更新: 2026-09-01
- 分类: learning
- 来源类型: paper
- 标签: llm-agent, benchmark, agent-harness, reasoning
- 质量评分: 4/5
- 抓取时间: 2026-09-03T04:29:22Z
中文导读
随着动态 agent harness 越来越多让 LLM 代理可以修改自身执行环境,代理需要理解组件生命周期CordisBench 提供 1,200 道题目,覆盖干走顺序状态预测不同顺序下的不变量与重配置选择评测发现三款效率型模型在低推理开销下只能处理小系统,随相关交互增加性能明显下降,特别是对打纳顺序的推理与状态预测可作为 harness 面向代理能力评估的新贝伦定位基准
为什么值得关注
随着动态 agent harness 越来越多让 LLM 代理可以修改自身执行环境,代理需要理解组件生命周期CordisBench 提供 1,200 道题目,覆盖干走顺序状态预测不同顺序下的不变量与重配置选择评测发现三款效率型模型在低推理开销下只能处理小系统...
关键信息
- 论文标题:CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
- 作者:Damien Sileo, Dimitri Kachler
- arXiv:https://arxiv.org/abs/2609.01600
- 发布时间:2026-09-01
- arXiv 分类:cs.CL, cs.AI
- 关联标签:llm-agent, benchmark, agent-harness, reasoning
English Abstract
Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting with programs executed against Cordis, a runtime that manages component dependencies and cleanup, and asks models to identify affected components, predict state after a specified teardown order, determine which conditions hold under all or some orders, and choose reconfigurations that succeed when executed. Across these tasks, we evaluate three efficiency-oriented models at low reasoning effort with 2, 4, 8, 16, 24, or 32 relevant interactions, using deterministic task-specific scoring. Models usually handle small systems well but grow less reliable as more interactions become relevant, especially when predicting final state and when reasoning across teardown orders. Additional inference effort recovers marked gains for some models. The cost is nontrivial: on our 16-interaction subset, GPT-5.6 Luna uses nearly 3,000 reasoning tokens per question at medium effort. For these controlled instances, that cost is avoidable: an independent finite reference semantics agrees with Cordis execution on every observation and action outcome used for scoring across all 528 executable questions.
English Summary
Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting with programs executed against Cordis, a runtime that manages component dependencies and cleanup, and asks models to identify affected components, predict state after a specified teardown order, determine which conditions hold under all or some orders, and choose reconfigurations that succeed when executed. Across these tasks, we evaluate three efficiency-oriented models at low reasoning effort with 2, 4, 8, 16, 24, or 32 relevant interactions, using deterministic task-specific scoring....
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。