Agent 与自动化 4.0 · 优秀 2026-08-26 · 论文

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

提出围绕持久跨交互规则记忆的自进化测试时防御:攻击得手时,框架把失败抽象为捕获结构性攻击包装(而非有害话题)的方法级规则并在后续输入复用;因为规则是方法级的,一条规则即可泛化整个攻击家族,标签空间随新包装出现而扩展机制完全依赖外部记忆与提示,不改参数,同时适用于开源权重与黑盒 API 模型在四类黑盒越狱家族与多个模型上大幅降低攻击成功率保住良性效用,在自适应复合包装攻击下保持稳健,且随记忆增长不推高过度拒答

打开原文回到归档

A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks

中文导读

提出围绕持久跨交互规则记忆的自进化测试时防御:攻击得手时,框架把失败抽象为捕获结构性攻击包装(而非有害话题)的方法级规则并在后续输入复用;因为规则是方法级的,一条规则即可泛化整个攻击家族,标签空间随新包装出现而扩展机制完全依赖外部记忆与提示,不改参数,同时适用于开源权重与黑盒 API 模型在四类黑盒越狱家族与多个模型上大幅降低攻击成功率保住良性效用,在自适应复合包装攻击下保持稳健,且随记忆增长不推高过度拒答

为什么值得关注

免参数更新的自进化防御:把越狱失败抽象成方法级规则存入跨交互记忆,一条规则泛化整个攻击家族;对部署黑盒 API 模型的团队,这是少数不改权重、不推高过度拒答也能持续压低攻击成功率的可行路线。

关键信息

  • 论文标题:A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks
  • 作者:Tongyan Hu, Bryan Hooi
  • arXiv:https://arxiv.org/abs/2608.26008
  • 发布时间:2026-08-26
  • arXiv 分类:cs.CR, cs.CL
  • 关联标签:safety、jailbreak、multi-agent、memory、arxiv

English Abstract

Large language models (LLMs) remain vulnerable to jailbreak attacks that exploit techniques such as role-playing, obfuscation, code transformation, and multi-step indirection to elicit harmful outputs. As jailbreak strategies keep emerging, defenses have proliferated in an ongoing cat-and-mouse game, yet most remain static: their safety behavior is fixed at deployment, so they cannot accumulate defensive experience or adapt to unseen strategies. We propose a self-evolving test-time defense built around a persistent, cross-interaction rule memory: when an attack succeeds, the framework abstracts that failure into a method-level rule capturing the structural attack wrapper rather than the harmful topic, and reuses it against future inputs. Because rules are method-level, one induced rule generalizes across an entire attack family, and the label space expands as novel wrappers appear. The mechanism operates entirely through external memory and prompting, with no parameter updates, and applies to both open-weight and black-box API models. We realize it as four cooperating modules, but the contribution is the memory-based adaptation mechanism, not the module decomposition. Across four black-box jailbreak families and multiple models, our method substantially reduces attack success rates while preserving benign utility, remains robust under an adaptive composite-wrapper attack, and does not increase over-refusal as the memory grows.

English Summary

A self-evolving test-time defense built on a persistent, cross-interaction rule memory: when a jailbreak succeeds, the framework abstracts the failure into a method-level rule capturing the structural attack wrapper rather than the harmful topic, then reuses it against future inputs. One induced rule generalizes across an entire attack family, and the label space expands as novel wrappers appear. The mechanism runs entirely through external memory and prompting with no parameter updates, applying to open-weight and black-box API models. Across four black-box jailbreak families and multiple models it substantially reduces attack success rates while preserving benign utility, stays robust under an adaptive composite-wrapper attack, and does not increase over-refusal as memory grows.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。