RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
- ID: f0ca5c22
- 原文链接: https://arxiv.org/abs/2608.27439
- PDF: https://arxiv.org/pdf/2608.27439v1
- 作者: Junjie Zhang, Hui Liu, Kecheng Chen, Xianbo Mo, Changsheng Chen, Haoliang Li
- 日期: 2026-08-27
- 更新: 2026-08-27
- 分类: learning
- 来源类型: paper
- 标签: security, red-teaming, agents, jailbreak, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-29T20:21:03+08:00
中文导读
RedEvoAgent(arXiv 2608.27439,cs.CR / cs.AI,2026-08-27)处理的是一个日益现实的风险场景:LLM 智能体越来越多地部署在产品级执行环境(execution harness)中,越狱一旦触发有害工具调用或持久状态变更,后果比单纯生成不安全文本严重得多。现有自动红队方法多依赖固定攻击;近期的 agentic 攻击者虽然会协调多个越狱工具、并借助基于轨迹的检索展现更强潜力,但检索偏置与工具功劳不清会复用误导性经验,完整轨迹还带来上下文开销并降低可解释性。RedEvoAgent 是黑盒红队智能体:把跨案例攻击轨迹蒸馏成简洁、人类可读的攻击技能;技能通过工具效果画像(tool-effectiveness profiling)与 Deciding-Tool Attribution 进行更新,并用验证棘轮(validation ratchet)只保留能提升验证集表现的更新。在多个基准、多个目标模型与多个目标执行环境上的实验显示,它优于固定攻击与 agentic 基线,提升工具使用效率,并可跨攻击者模型与目标执行环境迁移。
为什么值得关注
对做 agent 安全的团队,这条思路可以直接借鉴:把易被检索偏置污染的轨迹复用,换成"蒸馏后的可读攻击技能 + 只进不退的验证棘轮",攻击经验既可解释又能稳定积累。蒸馏出的攻击技能也适合作为外部评测集,用来检验自家 harness 的越狱防护、工具权限与状态变更审计设计——论文展示的跨执行环境迁移性正好说明这类威胁不绑定单一 harness。
原文(抓取存档·节选)
> Abstract (arXiv 2608.27439, v1 2026-08-27, cs.CR, cs.AI)
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profiling and Deciding-Tool Attribution for skill updates, and a validation ratchet that retains only updates improving validation performance. Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses.
Obsidian Notes
- 内容由
opencli arxiv paper 2608.27439 -f json拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。