AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- arXiv: 2608.20318
- Authors: Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
- Published: 2026-08-20; categories: cs.AI, cs.CL, cs.LG
Abstract
Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule improves the compute\mbox{-}capability exchange rate for every subsequent run, including the one that produces the next agent. Whether RSI is feasible therefore turns on whether an agent can design training algorithms. No benchmark isolates that ability: existing suites are won by collecting data or by tuning hyperparameters, and none tells a change to how a run is executed apart from a change to how the model learns. We present AI4AI\mbox{-}Bench, 10 frozen research repositories spanning 10 training algorithm families. In each task, an agent has 4 hours on one B300 to rewrite the training algorithm; its code is then rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent, against the repository's original algorithm under the same procedure. Because the 10 metrics are incommensurable, every task is mapped onto one scale on which $0$ is an uninformative model, $0.1$ is the algorithm the repository ships, and $1.0$ is the task optimum. Across 29 configurations of 6 systems on all 10 tasks the mean score is $0.166$, and the best system reaches $0.250$: even the strongest closes under a fifth of the distance between the algorithm that was already there and the optimum. The submissions show where that distance went: most never change how the model learns at all, and the minority that do average $0.226$ against $0.126$ for the rest. More reasoning effort mostly buys the willingness to go there, taking that minority from $8\%$ of submissions to $64\%$ and the mean score from $0.094$ to $0.196$. We release the task suite, the evaluators and every scored submission, so that the measurement can be repeated as these systems change.
Why it matters (AAIF scan)
递归自我改进 (RSI) 的可行性取决于 agent 能否设计训练算法AI4AI-Bench 给出首个隔离式基准:10 个冻结研究仓库覆盖 10 类训练算法族,agent 在单张 B300 上有 4 小时改写训练算法,改后代码从零重跑至多 12 小时,由隐藏的固定评估器与原算法同规程打分;所有任务映射到统一刻度(0=无信息模型,0.1=仓库自带算法,1.0=任务最优)29 种配置 6 个系统平均得分 0.166最好 0.250距最优连五分之一的路都没走完,为 RSI 能力划出了当前真实水位
Source: https://arxiv.org/abs/2608.20318
Captured: 2026-08-22 (AAIF content-fetcher)