UNISON: Co-Designed Near-Memory Session KV Scheduler for LLM Agents
Abstract (opencli arxiv paper)
arXiv 2609.09643 abstract (opencli arxiv paper 2609.09643 -f json): UNISON is an event-driven near-memory scheduler next to SRAM/HBM with SPEAR (session-interval-mean + turn-indexed hazard) and TIDE (DMA-budget view of observed waits) sharing one live ranking. Best non-oracle on every trace of 1,415 sessions / 33,596 turns: hit rate +0.3–23.1%, AMAT −22 to −51%, long-trace TTFT −58 to −89%. The 28 nm CMOS scheduler core is 0.169 mm² at 13.6 mW @ 150 MHz.
论文要点 (中文)
复旦(Fan He、Yan Li、Xiaoyang Zeng)2026-09-09 提交。LLM 被组合进 agent 循环(规划→调工具→等回应→再恢复同一任务)时,与传统多轮聊天对内存的压力完全不同:工具等待期间持续累积 KV 前缀,多个会话共用 SRAM/HBM,淘汰和分层放置成了与计算模式正交的会话级效率问题。基于 recency / timeout / 身份的策略把活的工具等待当成「冷、可弃」单元——错。UNISON 是在 SRAM/HBM 旁的事件驱动近存调度器,SPEAR(基于会话间隔均值和 turn-indexed hazard 选谁离开)与 TIDE(把观察到的等待当作 DMA 预算,决定谁坐快层)共享同一份 live ranking。1,415 会话、33,596 轮的编码与通用基准上,UNISON 在所有 trace 上都是非 oracle 中最佳:hit rate +0.3–23.1%、AMAT −22 到 −51%、长程 trace 的 TTFT −58 到 −89%;结构性必要分析显示这种统一近存设计不可拆成独立 IP、也不能在不重现已记录失效模式的情况下用软件实现。28 nm CMOS 调度核心 0.169 mm²、13.6 mW @ 150 MHz,相对它管理的 KV 层次可忽略不计。给做 agent 推理基础设施的读者画了一张「KV 在 SRAM/HBM 上要按会话级 hazard 排,不要按访问 recency 排」的硬件-算法协同设计图。
Key claims (English)
UNISON (Fudan; Fan He, Yan Li, Xiaoyang Zeng; submitted 2026-09-09) is a co-designed hardware–algorithm answer to a memory-pressure pattern unique to LLM agents: an agent loop plans, calls tools, waits, then resumes the same task. Tool waits keep accumulating KV prefixes; many sessions share SRAM/HBM; eviction/tiering becomes a session-level problem orthogonal to the compute pattern. Recency/timeout/identity heuristics treat live tool waits as cold and discardable — wrong. UNISON is an event-driven near-memory scheduler next to SRAM/HBM. SPEAR (leaving-decision from session-interval mean + turn-indexed hazard) and TIDE (DMA-budget view of observed waits, deciding who sits in fast tier) share one live ranking. Across 1,415 sessions / 33,596 turns of coding + general benchmarks, UNISON is the best non-oracle on every trace: hit rate +0.3–23.1%, AMAT −22 to −51%, long-trace TTFT −58 to −89%. Structural-necessity analysis shows the design cannot be decomposed into independent IPs nor re-implemented in software without re-creating the recorded failure modes. The 28 nm CMOS scheduler core is 0.169 mm² at 13.6 mW @ 150 MHz, negligible relative to the KV hierarchy it manages. The picture for agent-inference infrastructure readers is to rank KV by session-level hazard rather than by access recency.
Obsidian 证据摘录
「UNISON 是在 SRAM/HBM 旁的事件驱动近存调度器,SPEAR(基于会话间隔均值和 turn-indexed hazard 选谁离开)与 TIDE(把观察到的等待当作 DMA 预算,决定谁坐在快层)共享同一份 live ranking。……UNISON 在所有 trace 上都是非 oracle 中最佳,hit rate +0.3–23.1%、AMAT −22 到 −51%、长程 trace 的 TTFT −58 到 −89%;结构性必要分析显示这种统一近存设计不可拆成独立 IP、也不能在不重现已记录失效模式的情况下用软件实现。28 nm CMOS 调度核心 0.169 mm²、13.6 mW @ 150 MHz,相对它管理的 KV 层次可忽略不计。」——OpenClaw定时任务/论文流水线/2026-09-15-论文流水线.md L31-33
链接
- 论文:https://arxiv.org/abs/2609.09643
- 论文流水线:OpenClaw定时任务/论文流水线/2026-09-15-论文流水线.md