CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
- ID: 56c1bd50
- 原文链接: https://arxiv.org/abs/2608.21278
- PDF: https://arxiv.org/pdf/2608.21278v1
- 作者: Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo
- 日期: 2026-08-21
- 更新: 2026-08-21
- 分类: models
- 来源类型: paper
- 标签: llm-safety, lora-adapter, alignment, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-25T05:20:39+00:00
中文导读
安全对齐的老难题是全局微调同时伤及良性输入的用户体验:本文提出 CLEAR,用一个轻量隐状态门控连续地控制安全 LoRA 适配器的激活强度,只在判定需要时抵御有害请求,冻结的主干权重不动在 Llama-3-8B-Instruct 上,CLEAR 在 HarmBench 提升稳健性的同时,把全局 SFT/LoRA 安全微调带来的良性性能下降降了下来(摘要报告 ASR 从 32.x% 起降)对需要兼顾安全与产品体验的部署方是一条可落地的路径
为什么值得关注
用隐状态门控按输入调节安全 LoRA 的激活强度,安全与可用性不再二选一 — Per the abstract, CLEAR reports HarmBench ASR dropping from 32.3% to 0.5% on Llama-3-8B-Instruct while retaining most base utility, and up to +7.1 points GSM8K accuracy versus globally applied SFT/LoRA. For deployments that need both safety robustness and benign-prompt quality, gated continuous adapter routing is a directly implementable mechanism that leaves the frozen backbone untouched.
关键信息
- 论文标题:CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
- 作者:Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo
- arXiv:https://arxiv.org/abs/2608.21278
- 发布时间:2026-08-21
- arXiv 分类:cs.AI
- 关联标签:llm-safety, lora-adapter, alignment, arxiv
English Abstract
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign inputs. We propose Continuous LatEnt Adapter Routing (CLEAR), a conditional safety adaptation framework that uses a lightweight hidden-state gate to continuously control the activation strength of a safety low-rank adapter. CLEAR aims to reduce harmful completions while avoiding unnecessary changes to the frozen backbone that could degrade performance on benign prompts. Experiments on widely used safety and utility benchmarks show that CLEAR improves robustness on HarmBench while reducing the utility degradation observed with globally applied safety tuning such as SFT or standard low-rank adaptation (LoRA). On Llama-3-8B-Instruct, CLEAR reduces HarmBench ASR from 32.3% to 0.5%, while retaining most of the base model's utility and achieving up to 7.1 percentage points higher GSM8K accuracy than globally applied SFT or LoRA. These results suggest that CLEAR is a promising mechanism for improving the safety--utility trade-off in LLM alignment.
English Summary
CLEAR attaches a lightweight hidden-state gate that continuously modulates the activation strength of a safety low-rank adapter, leaving the frozen backbone untouched on benign prompts. On Llama-3-8B-Instruct it improves HarmBench robustness while reducing the utility degradation that globally applied safety tuning (SFT/LoRA) causes.
Obsidian Notes
- 内容由
opencli arxiv paper拉取 arXiv 元数据与摘要生成。 - 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。