CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
- ID: ba944e18
- 原文链接: https://arxiv.org/abs/2609.26779
- PDF: https://arxiv.org/pdf/2609.26779v1
- 作者: Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers
- 发布日期: 2026-09-22
- 更新日期: 2026-09-22
- 分类: cs.AI
- 来源类型: paper
- 标签: arxiv, context-engineering, coding-agent, long-horizon, paper
- 质量评分: 4/5
- 抓取时间: 2026-09-24T04:23:22Z
中文导读
编码 agent 处理复杂任务需要跨会话的百万 token 级上下文,压缩不可避免CliffCompaction 是一种 autocompaction 技术:在受限上下文下将成本最多降 50%,Terminal-Bench 性能持平或更好,并在 KernelBench 上取得 SOTA关键设计是保真压缩只截断或丢弃内容,绝不改写;且永不压缩压缩产物,每轮只处理原始内容并丢弃旧压缩输出,防止上下文漂移累积这支撑了超百万 token 会话上的持续学习:KernelBench 上 CUDA kernel 加速 200 步达 2.23x400 步达 3.58x,超过专用搜索算法和训练型 agent已开源 scaffold 无关的 API-proxy 实现
为什么值得关注
只删不改永不二次压缩的 autocompaction:长程编码 agent 成本降 50%,百万 token 会话上 KernelBench 加速 3.58x
关键信息
- 论文标题:CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
- 作者:Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers
- arXiv:https://arxiv.org/abs/2609.26779
- 发布时间:2026-09-22
- 更新时间:2026-09-22
- arXiv 分类:cs.AI, cs.LG, cs.SE
- 可选注释:N/A
- 关联标签:arxiv, context-engineering, coding-agent, long-horizon, paper
English Abstract
Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of $2.23\times$ after 200 steps and $3.58\times$ after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.
English Summary
Agents working on complex problems need millions of tokens of context, forcing compaction across sessions. CliffCompaction is an autocompaction technique that cuts cost by up to 50% under a bounded context while maintaining or improving Terminal-Bench performance and achieving state-of-the-art results on KernelBench. It keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it, and never compacts a compaction: each pass operates only on original content and prior compacted output is discarded, preventing context drift from accumulating. This sustains continual learning over sessions exceeding a million tokens, reaching 2.23x CUDA kernel speedups after 200 steps and 3.58x after 400 steps on KernelBench. A scaffold-agnostic API-proxy implementation is open-sourced.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
- 本页面属于 content-fetcher 能动下的高分筻补东生成,不修改原条目字段,仅产出可读的中英双语页面。