Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
- ID: 37687d1d
- Original: https://arxiv.org/abs/2608.09900
- PDF: https://arxiv.org/pdf/2608.09900v1
- Authors: Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde, Gonzalo Martínez, Pedro Reviriego
- Published: 2026-08-10
- Updated: 2026-08-10
- arXiv Categories: cs.CL
- AAIF Category: models
- Source Type: paper
- Tags: robustness, decoding, evaluation, stress-test, llm
- Quality Score: 4/5
- Fetched At: 2026-08-12T04:29:18.480369+00:00
Chinese Guide
提出 Decoding-Level Taboo:一种零提示(zero-prompt)的 LLM 健壮性诊断压力测试,直接在 logit 空间运行时干预,动态遮蔽首要候选词迫使模型"绕道表达",以打破模型在名义生成路径上形成的"能力假象",评测多个开源模型族发现 off-path 健壮性同时受参数规模与指令对齐训练影响,且随规模/对齐增强而提升,可作为合成数据生成运行时安全护栏审计的通用原语
Why It Matters
提出 Decoding-Level Taboo:一种零提示(zero-prompt)的 LLM 健壮性诊断压力测试,直接在 logit 空间运行时干预,动态遮蔽首要候选词迫使模型"绕道表达",以打破模型在名义生成路径上形成的"能力假象"...
This content page is grounded in the existing AAIF entry plus arXiv metadata fetched with opencli arxiv paper. It does not add experimental claims beyond the arXiv abstract and entry summary.
Key Metadata
- Paper title: Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
- Authors: Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde, Gonzalo Martínez, Pedro Reviriego
- arXiv: https://arxiv.org/abs/2608.09900
- PDF: https://arxiv.org/pdf/2608.09900v1
- Published: 2026-08-10
- Updated: 2026-08-10
- arXiv categories: cs.CL
- Related AAIF tags: robustness, decoding, evaluation, stress-test, llm
English Abstract
Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safety guardrails, and structural constraints continuously force models off this nominal path, driving a divergence between benchmark scores and deployment performance. To address this issue, we introduce Decoding-Level Taboo, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths. By dynamically masking primary candidate tokens at word boundaries, Taboo forces machine circumlocution. Evaluating Taboo across several open-weight model families reveals that off-path robustness is heavily influenced by both parameter scale and post-training instruction alignment, with robustness generally improving with model size and alignment. Beyond the results presented in this paper, Taboo provides a novel primitive for generating diverse synthetic datasets, stress-testing runtime safety guardrails, and auditing model reliability prior to real-world deployment.
English Summary
Decoding-Level Taboo intervenes directly in logit space at runtime by dynamically masking primary candidate tokens at word boundaries, forcing models off their nominal generation paths and exposing the gap between benchmark scores and deployment robustness. Evaluating Taboo across several open-weight model families shows off-path robustness is heavily influenced by parameter scale and post-training instruction alignment, generally improving with both. Beyond diagnostics, Taboo serves as a primitive for synthetic data generation, runtime safety guardrail stress-testing, and pre-deployment reliability auditing.
Obsidian Notes
- Generated by the AAIF
content-fetchercron mode from arXiv metadata and the existing AAIF entry. - Chinese guide and relevance notes are anchored to the entry summary, arXiv abstract, authors, dates, categories, and links.
- Canonical content path is
content/{entry_id}.mdunder the repository root;openclaw/content/is intentionally avoided.