Agent 与自动化 4.0 · 优秀 2025-12-24 · 论文

A Plan Reuse Mechanism for LLM-Driven Agent

LLM 驱动的个人助手生成计划的延迟可达数十秒,真实数据分析显示约 30% 的请求与历史请求相同或相似,存在复用空间论文提出计划复用机制,缓存并复用已生成计划以降低交互延迟收录理由:面向消费级 agent 的延迟优化路线,与 KV cache/记忆方向互补

打开原文回到归档

A Plan Reuse Mechanism for LLM-Driven Agent

  • ID: 97fc809c
  • 原文链接: https://arxiv.org/abs/2512.21309
  • PDF: https://arxiv.org/pdf/2512.21309
  • 作者: Guopeng Li, Ruiqi Wu, Haisheng Tan
  • 日期: 2025-12-24
  • 更新: 2025-12-25
  • 分类: agents
  • 来源类型: paper (arxiv)
  • 标签: plan-reuse, llm-agents, latency-optimization, personal-assistant, arxiv
  • 质量评分: 4/5
  • 抓取时间: 2026-08-15T12:25:21Z

中文导读

LLM 驱动的个人助手(如小爱、Blue Heart V)生成计划延迟可达数十秒,直接拖垮交互体验。作者对真实数据集的分析显示约 30% 的请求与历史请求相同或相似,存在计划复用空间;难点在于原始请求文本的相似度难以直接定义——自然语言表达多样、计划文本非结构化。AgentReuse 利用请求语义间的相似与差异,用意图分类(intent classification)评估请求相似度并启用计划复用。真实数据集实验:93% 有效计划复用率,请求相似度评估 F1 0.9718 / 准确率 0.9459,相比无复用基线延迟降低 93.12%。

为什么值得关注

面向消费级 agent 的延迟优化路线:不靠更快的模型或更长的 KV cache,而是靠“请求级缓存命中”。30% 重复请求 + 93% 复用率 + 93% 延迟降幅这组数字,对任何做端侧/消费级 agent 交互延迟优化的人都是直接可对标的目标区间。把复用问题转化为意图分类问题这一步,也避开了文本相似度直接度量的坑。

关键信息

  • 论文标题:A Plan Reuse Mechanism for LLM-Driven Agent(系统名 AgentReuse)
  • 问题:LLM 生成计划的延迟可达数十秒;约 30% 请求与历史请求相同或相似
  • 难点:请求文本相似度难以直接定义;自然语言表达多样;计划文本非结构化
  • 方法:基于请求语义的意图分类评估相似度,复用已生成计划
  • 关键数字:93% 有效复用率;F1 0.9718;准确率 0.9459;延迟降低 93.12%
  • 场景:小爱、Blue Heart V 等 LLM 驱动个人助手(含 IoT 设备管理)
  • arXiv 分类:cs.MA

English Abstract

Integrating large language models (LLMs) into personal assistants, like Xiao Ai and Blue Heart V, effectively enhances their ability to interact with humans, solve complex tasks, and manage IoT devices. Such assistants are also termed LLM-driven agents. Upon receiving user requests, the LLM-driven agent generates plans using an LLM, executes these plans through various tools, and then returns the response to the user. During this process, the latency for generating a plan with an LLM can reach tens of seconds, significantly degrading user experience. Real-world dataset analysis shows that about 30% of the requests received by LLM-driven agents are identical or similar, which allows the reuse of previously generated plans to reduce latency. However, it is difficult to accurately define the similarity between the request texts received by the LLM-driven agent through directly evaluating the original request texts. Moreover, the diverse expressions of natural language and the unstructured format of plan texts make implementing plan reuse challenging. To address these issues, we present and implement a plan reuse mechanism for LLM-driven agents called AgentReuse. AgentReuse leverages the similarities and differences among requests' semantics and uses intent classification to evaluate the similarities between requests and enable the reuse of plans. Experimental results based on a real-world dataset demonstrate that AgentReuse achieves a 93% effective plan reuse rate, an F1 score of 0.9718, and an accuracy of 0.9459 in evaluating request similarities, reducing latency by 93.12% compared with baselines without using the reuse mechanism.

English Summary

LLM-driven assistants can take tens of seconds to generate a plan, degrading user experience. Real-world dataset analysis shows about 30% of requests are identical or similar to prior ones, enabling plan reuse: the paper proposes a mechanism that reuses previously generated plans to cut interaction latency in LLM-driven agents.

Obsidian Notes

  • 内容由 opencli arxiv paper 2512.21309 -f json 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断锚定在 arXiv 摘要与条目既有 summary 上;未补充摘要之外的实验细节。
  • 条目收录日期:2026-08-15;语言:both。