Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
- ID: f9084c0b
- 原文链接: https://arxiv.org/abs/2608.24188
- PDF: https://arxiv.org/pdf/2608.24188v1
- 作者: Jiayu Shi, Luzhuo Chen
- 日期: 2026-08-25
- 更新: 2026-08-25
- 分类: coding
- 来源类型: paper
- 标签: coding-agent, context-compression, lora, token-efficiency, arxiv
- 质量评分: 4/5
中文导读
Coding agent 每轮重发整块文件读取与工具输出,这部分上下文撑起 token 支出大头;通用压缩器按散文训练,在代码上会改写标识符、丢掉待编辑的精确行。Paritok-4B 是 4B LoRA 抽取式压缩器,两条原则:只选不改——输出的标识符、路径、数字 96.0% 直接来自输入,held-out 上保持 96.2%;带任务意图选行——保留行比删除行的意图相关度平均高 0.067(95% CI [+0.056, +0.078])。训练由 gpt-4.1-mini 教师从 67,074 条真实 OpenHands 轨迹蒸馏出 40,606 个校验样例,底座 Qwen3-4B。SWE-bench Lite 全部 300 实例上压缩到原来的 25.7%(gpt-4.1-mini 为 50.2%,gpt-5 为 61.9%),保留未压缩单次解题质量的 86.5%;真实带行号输入下 27.8% 压缩、89.3% 保留,McNemar p=0.079 说明该压缩幅度下解题率无统计显著下降。成品是 264 MB adapter,24 GB 单卡自托管,无按 token 压缩费——文中算账:按牌价拿 gpt-5 当压缩器净亏。权重与数据 Apache 2.0 开源。
为什么值得关注
Paritok-4B:抽取式+意图条件化把 agent 上下文压到 25.7% 仍保 86.5% 解题质量,264 MB adapter 单卡自托管免压缩费
English Abstract
Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context dominates their token bill. General-purpose prompt compressors are trained on prose and suit code poorly: they paraphrase identifiers and drop the exact spans an agent needs to edit. We present Paritok-4B, a 4B LoRA compressor for coding-agent trajectories built on two commitments. It is extractive: it selects spans rather than rewriting them, and 96.0% of the identifiers, paths, and numbers it emits already appear in its input, holding at 96.2% on held-out SWE-bench Lite output. It is intent-conditioned: told the agent's current task, it acts chiefly inside a retained segment, selecting which lines survive (retained lines are +0.067 more intent-relevant than removed ones, paired 95% CI [+0.056, +0.078]) rather than changing how much is retained. We distil a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories into 40,606 validated examples and fine-tune Qwen3-4B. On all 300 SWE-bench Lite instances, Paritok-4B compresses agent context to 25.7% of its size, 2.0x harder than a gpt-4.1-mini compressor (50.2%) and 2.4x harder than gpt-5 (61.9%), while retaining 86.5% of uncompressed single-shot solve quality. Fed the cat -n line-numbered input real agents produce, it compresses slightly less (27.8%) and retains more (89.3%); there the paired test is informative, with 30 instances solved only uncompressed and 17 only compressed, an exact McNemar p=0.079, so at this sample size compressing context to roughly a quarter of its size does not significantly reduce the solve rate. The model is a 264 MB adapter that self-hosts on one 24 GB GPU with no per-token compressor fee, which at list prices decides the economics: gpt-5 as a compressor is net-negative, costing more than the downstream tokens it saves. Weights, data, and evaluation scripts are open (Apache 2.0).
Obsidian Notes
- Metadata and abstract fetched via
opencli arxiv paper 2608.24188 -f json(2026-08-27); response parsed list-or-dict tolerant. - 数字(25.7%/86.5%/96.0%/40,606 样例/McNemar p=0.079)均直接取自摘要。
- 中文导读与价值判断锚定在论文摘要上,未补充摘要之外的实验细节。