Muscle Memory for Agents: Compile not Merely Retrieve
Source: https://arxiv.org/abs/2608.08995
Author: Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang, Tanya Dixit
Original date: 2026-08-10
Added by: AAIF daily-intake-evening 2026-08-13
摘要
论文提出 Muscle Memory:不要把反复出现的用户意图只作为文本记忆检索,而是编译成带触发器的 specialist agent。Harvest→Analyze→Augment→Evaluate 四阶段从对话历史中分离 behavioral pattern 与 task pattern,并在 90 个 held-out 场景中显示 specialist 触发时 88.9% 胜率、个性化收益 +2.05。
English Summary
Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpose orchestrator interpret what to do. This paper argues that the pattern is the wrong default for personalization. We position Muscle Memory - the practice of compiling recurring user intent into purpose-built specialist agents - as a distinct memory paradigm from retrieval, and we argue that compilation is a better fit for the workloads where current assistants impose a multi-turn tax on their users: making them repeatedly correct format, depth, and scope to obtain a domain-appropriate answer. We support the position with a reference implementation and empirical evidence. The implementation is a four-phase pipeline (Harvest $\rightarrow$ Analyze $\rightarrow$ Augment $\rightarrow$ Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable compiled specialists with two-stage trigger matching. On 90 held-out scenarios across five user personas, the augmented assistant wins 32 of 36 cases where a specialist fires, an 88.9% win rate, with a +2.05 personalization gain and only a $-0.28$ accuracy cost on a 1-4 scale. We discuss why compilation is better suited than retrieval in this regime, what the result implies for the broader memory design space, and what open problems remain.
入库理由
- quality_score: 5
- category: agents
- tags: agent-memory, personalization, specialist-agents, compile-not-retrieve
- one_liner: Agent 记忆不只应 retrieve;高频意图可以 compile 成可触发的 specialist agent。
Obsidian evidence excerpt
rXiv:2608.11079。把 self-evolving agent 累积的 skill 当成"typed contract"——名字/描述/工作流/工具契约/输出字段/例外规则——用 typed minimum description-length 目标一次性压缩:重复规则 state once at scope,重复动作序列 factor into shared procedure,例外保留为 explicit exceptions。Zip-on-Write 模式支持 incremental 演化不重放任务。压缩率高、保留 unique rare rules by construction。
来源:https://arxiv.org/abs/2608.11079
2. **Muscle Memory for Agents: Compile not Merely Retrieve** — arXiv:2608.08995。主张把"反复出现的用户意图"编译成 purpose-built specialist agent,而不是 retrieve-then-orchestrate。Harvest→Analyze→Augment→Evaluate 四阶段管线,从对话历史中分别挖出 behavioral pattern 和 task pattern,发出的 specialist 配 two-stage trigger matching。90 held-out scenario 上 specialist 触发时 88.9% 胜率,+2.05 personalisation gain,accuracy 损失仅 -0.28(1-4 scale)。
来源:https://arxiv.org/abs/2608.08995
3. **GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning** — arXiv:2608.10494。Training-free self-evolving 框架,把完成的轨迹压成结构化 nonparametric execution state——Workflow Graph Memory(全局操作顺序)+ Action-Level Experiences(局部纠错)+ Adapted Skill SOP(程序与数据约束)。执行、蒸馏、复用三段循环里 backbone LLM 不变。多个 geospatial benchmark 上同时拉高 task accuracy 和 tool-use trajectory quality。
来源:https://arxiv.org/abs/2608.10494
4. **MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows** — arXiv:2608.10509。把 agent、source、memory、claim、action 全部建模成 typed execution graph:lineage tracing + permission-ineligible record exclusion + semantic similarity × multiplicative path trust reranking + risk-sensitive action gate。2,700 合成任务上 94.96% task success / 72.70% exact decision accuracy / 90.22% clean。把 provenance 当成 operational control signal,而不是事后审计。
来源:https://arxiv.org/abs/2608.10509
5. **ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization** — arXiv:2608.11045。Post-training 量化新路线:训 conditional diffusion model 生成 low-bit weight 的连续 reconstruction,给处于 quantization interval 中点附近的 weight 提供 rounding direction guidance。Tolerance metric 决定哪些 weight 走 diffusion reconstruction、哪些走 RTN。Calibration-free、offline、inference 无额外开销,小 LLM 上 3-bit / 4-bit 都稳定超过 RTN。
来源:https://arxiv.org/abs/2608.11045
6. **Neural Introspectio
arXiv metadata / abstract
- arXiv id: 2608.08995
- authors: Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang, Tanya Dixit
- published: 2026-08-10
- updated: 2026-08-10
- categories: cs.MA
- PDF: https://arxiv.org/pdf/2608.08995v1
Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpose orchestrator interpret what to do. This paper argues that the pattern is the wrong default for personalization. We position Muscle Memory - the practice of compiling recurring user intent into purpose-built specialist agents - as a distinct memory paradigm from retrieval, and we argue that compilation is a better fit for the workloads where current assistants impose a multi-turn tax on their users: making them repeatedly correct format, depth, and scope to obtain a domain-appropriate answer. We support the position with a reference implementation and empirical evidence. The implementation is a four-phase pipeline (Harvest $\rightarrow$ Analyze $\rightarrow$ Augment $\rightarrow$ Evaluate) that mines conversational history, separates behavioral from task patterns, and emits quality-gated executable compiled specialists with two-stage trigger matching. On 90 held-out scenarios across five user personas, the augmented assistant wins 32 of 36 cases where a specialist fires, an 88.9% win rate, with a +2.05 personalization gain and only a $-0.28$ accuracy cost on a 1-4 scale. We discuss why compilation is better suited than retrieval in this regime, what the result implies for the broader memory design space, and what open problems remain.