Agent 与自动化 5.0 · 必读 2026-08-05 · 论文

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

Argus 把长程推理视为可持久的 agentic runtime,由 ManagerPlannerEngineerReviewer 在持久项目状态上执行有边界任务摘要强调它不改模型权重,而是通过 memoriesskillsproceduresverifiers和路由决策的审核入库来自我演化论文报告在 GPT-5.5 七个基准中 SWE-Bench Pro 约 78%,并在验证门控后降低求解 token 和工作时间

打开原文回到归档

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

  • ID: 46235359
  • Original URL: https://arxiv.org/abs/2608.05144
  • PDF: https://arxiv.org/pdf/2608.05144v1
  • Author(s): Boxiu Li, Zimo Wen, Yijia Fan, Junxiang Lei, Sufeng Guo, Jiaao Wu, Ruize Tang, Mukai Li, Yifei Shen, Xiaoyu Chen, Wanbo Zhang, Runjing Gu, Yifei Gao, Yuheng Wu, Xuyao Huang, Zelong Zhao, Jiachen Zhang, Shibo Hu, Hangxi Guo, Yilin Chen, Yuzhe Zhang, Fan Yang, Chuan Wen, Xian Zhang, Xuanhe Zhou, Zhijie Deng
  • Date: 2026-08-05
  • Category: cs.AI
  • Source type: paper
  • Tags: agent-runtime, long-horizon-reasoning, self-evolving-agents, verification, swe-bench
  • Quality score: 5/5
  • Fetched at: 2026-08-07T04:20:20+00:00
  • Obsidian evidence: OpenCLI arXiv metadata backfill

中文导读

Argus 把长程推理视为可持久的 agentic runtime,由 ManagerPlannerEngineerReviewer 在持久项目状态上执行有边界任务摘要强调它不改模型权重,而是通过 memoriesskillsproceduresverifiers和路由决策的审核入库来自我演化论文报告在 GPT-5.5 七个基准中 SWE-Bench Pro 约 78%,并在验证门控后降低求解 token 和工作时间

为什么值得关注

This fills a high-score (5/5) AAIF content gap around agent-runtime, long-horizon-reasoning, self-evolving-agents, verification, with the abstract giving enough grounded detail for follow-up reading and comparison.

English Summary

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points....

原文摘要 / Source Excerpt

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

  • arXiv: https://arxiv.org/abs/2608.05144
  • PDF: https://arxiv.org/pdf/2608.05144v1
  • Authors: Boxiu Li, Zimo Wen, Yijia Fan, Junxiang Lei, Sufeng Guo, Jiaao Wu, Ruize Tang, Mukai Li, Yifei Shen, Xiaoyu Chen, Wanbo Zhang, Runjing Gu, Yifei Gao, Yuheng Wu, Xuyao Huang, Zelong Zhao, Jiachen Zhang, Shibo Hu, Hangxi Guo, Yilin Chen, Yuzhe Zhang, Fan Yang, Chuan Wen, Xian Zhang, Xuanhe Zhou, Zhijie Deng
  • Published: 2026-08-05
  • Updated: 2026-08-05
  • Categories: cs.AI

Abstract

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.