CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
摘要中文导览(来自条目评分时的双语摘要,基于论文摘要原文提炼):
- 推理时扩展通常依赖重复生成或外部验证,CritICL 的核心洞察是同一模型家族内 LLM 失败模式在不同规模间呈现结构化规律它不把失败当作废料,而是从小模型提取失败模式,以批判式 in-context 示例注入大模型推理过程提供两个变体:CritICL-dynamic 按输入自适应预测可能的失败模式并检索对应批判;CritICL-static 用全局失败模式画像给出稳定引导实验显示其持续优于标准 ICL,达到与 test-time scaling 相当或更好的表现,同时生成次数与 token 成本显著更低代码已开源
论文信息
- arXiv ID: 2608.27455
- 作者: Yufan Wu, Yinghui He, Zhengyi Hu, Lang Wei, Ruichen Li 等(7 人)
- 发表: 2026-08-27(更新:2026-08-27)
- 分类: cs.CL
- 备注: https://github.com/umwyf/CRITICL
- 原文链接: https://arxiv.org/abs/2608.27455
- PDF: https://arxiv.org/pdf/2608.27455v1
- 标签:
inference-time-scalingweak-to-strongfailure-modesin-context-learningreasoning
Abstract(原文)
Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes exhibit structured patterns across model scales within the same family. Instead of treating failures as undesirable outputs, CritICL leverages them as a source of guidance. Specifically, we utilize failure modes derived from weaker models and incorporate them into inference through critique-based in-context examples. We propose two variants: CritICL-dynamic, which adaptively predicts input-specific failure modes and retrieves critiques, and CritICL-static, which uses a global failure mode profile to provide stable guidance. Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost. Code available at: https://github.com/umwyf/CRITICL
核心要点(英文摘要的中文提炼)
- 论文题为 CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes,发表于 arXiv(cs.CL,2026-08-27 提交,2026-08-27 更新)。
- 上述中文导览对应摘要中声明的贡献、实验设置与结论,未引入摘要之外的事实。
- 建议阅读顺序:先看 Abstract 原文核对该论文的动机与方法声明,再按需下载 PDF 深入实验细节。