Agent 与自动化 4.0 · 优秀 2026-09-14 · 论文

The Router Within: Eliciting Native Skill Routing from a Frozen LLM

arXiv 2609.15982(cs.LG/AI/CL,2026-09-14,Chen/Wang/Chen/Li/Huang)指出当前 agent 路由方法要么把所有 skill metadata 塞进上下文(分散注意力上限大小),要么外置检索(脱离模型能力)本文用 Gavel(Glance And Verdict)从冻结 LLM 的中间层前向特征中读出原生路由信号:两线性映射 + per-skill bankGlance打分,shortlist 的 skill 再续做前向Verdict读模型自身 likelihood 与 yes/no 判断,二者以 product-of-experts 融合只训练两个映射...

打开原文回到归档

内置路由器:从冻结 LLM 本身抽取原生技能路由能力

Source: arXiv:2609.15982 · Category: cs.LG · Authors: Ruishuo Chen, Xun Wang, Yu Chen, Zhuoran Li, Longbo Huang

TL;DR

技能(skills)将 LLM 智能体拓展出参数知识以外,其收益取决于能否选对技能。部署中的路由框架依赖预先技能元数据装入上下文,会拉散智能体注意力、限制能装的技能库规模;检索流水线把选择拆出上下文,却也把选择拆出了智能体能力。本文证明:冻结的智能体 LLM 本身在其前向传递中已经携带了路由信号,只需两个线性映射就能读出该信号,上下文中无需任何技能文本。Gavel(Glance And Verdict from a frozen LLM)分两步:Glance 在初装时以一次前向构建每技能的紧紧的线性铺,接着将任务与每个技能中层状态投射过两个可训练矩阵(也是全部训练参数)后评分整个技能库;Verdict 再恢复入选技能的前向传递,读取模型本身的似然与 yes/no 判断,与 Glance 以专家乘积方式融合。只训一次,模型零样本迁移到三个公开基准与 SkillTraj(本文发布的 372 条仿真智能体轨迹基准)。在 Qwen3-32B 上超越加 1.2B–16B 外部参数的进阶披露与检索-重排流水线:文本任务高 13.4 点,位于交互中途需要技能的场景高 21.9 点。路由准确率随背骨模型升级而提升;同一个 32B 在 bash 智能体中在 Skill-Use 上命中正确技能的次数还多于跑在 Codex 里的很大型号前沿模型。

为什么重要

能库越大、上下文越短,智能体能越需要一个“原生路由能力”:不装技能文本、不加检索器、只读冻结 LLM 本身的中层状态。这个路子可能是越大能能库 / 多智能体的现实路径。

原文摘要(English)

Skills extend an LLM agent beyond its parametric knowledge, and the gain they promise rests on picking the right one. Deployed harnesses route by preloading every skill's metadata into the context, which disperses the agent's attention and caps the library size. Retrieval pipelines move the selection out of the context, but also out of the agent's capability. We show that the frozen agent LLM already carries the routing signal in its own forward passes, and that two linear maps suffice to read it out with no skill text in the context. Gavel (Glance And Verdict from a frozen LLM) reads it in two steps. A glance projects the task's and each skill's mid-layer states through the two maps, the only parameters trained, and scores the full library against compact per-skill banks that one forward pass builds at installation. A verdict then resumes the shortlisted skills' forward passes and reads the model's own likelihood and yes/no judgment, fused with the glance as a product of experts. Trained once, Gavel transfers zero-shot to three public benchmarks and SkillTraj, our new benchmark of 372 simulated agent trajectories. On Qwen3-32B it outperforms progressive disclosure and progressive retrieve-and-rerank pipelines that add 1.2B to 16B external parameters, by up to 13.4 points on written tasks and up to 21.9 when the need for a skill arises mid-rollout. Routing accuracy improves as the backbone does, and in a bash-agent harness the same 32B triggers the correct skill on Skill-Use more often than far larger frontier models running in Codex.

基本信息

| 项 | 值 | |------|------| | 论文 ID | 2609.15982 | | 主分类 | cs.LG | | 发表 | 2026-09-14 | | 更新 | 2026-09-14 | | 作者 | Ruishuo Chen, Xun Wang, Yu Chen, Zhuoran Li, Longbo Huang | | PDF | 2609.15982 | | abs | https://arxiv.org/abs/2609.15982 |

参考