研究与学习 4.0 · 优秀 2026-08-05 · 论文

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Skill-Native LLMs 把长程推理中的跨技能切换抽象成 Skill Entropy,用来衡量模型从一种推理技能转向另一种技能的难度摘要介绍了基于 558 个技能9 个领域的 Skill^2-Bench,并在 8 个前沿模型和 4 个开源模型上观察到高熑任务准确率下降它还将 skill entropy 转成 RL 训练奖励,报告 Qwen3-4B-Instruct 分数从 34.4% 提升到 68.4%

打开原文回到归档

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

  • ID: 88153c92
  • Original URL: https://arxiv.org/abs/2608.05139
  • PDF: https://arxiv.org/pdf/2608.05139v1
  • Author(s): Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora
  • Date: 2026-08-05
  • Category: cs.CL, cs.LG
  • Source type: paper
  • Tags: skill-entropy, long-horizon-reasoning, benchmark, reinforcement-learning, skill-switching
  • Quality score: 4/5
  • Fetched at: 2026-08-07T04:20:20+00:00
  • Obsidian evidence: OpenCLI arXiv metadata backfill

中文导读

Skill-Native LLMs 把长程推理中的跨技能切换抽象成 Skill Entropy,用来衡量模型从一种推理技能转向另一种技能的难度摘要介绍了基于 558 个技能9 个领域的 Skill^2-Bench,并在 8 个前沿模型和 4 个开源模型上观察到高熑任务准确率下降它还将 skill entropy 转成 RL 训练奖励,报告 Qwen3-4B-Instruct 分数从 34.4% 提升到 68.4%

为什么值得关注

This fills a high-score (4/5) AAIF content gap around skill-entropy, long-horizon-reasoning, benchmark, reinforcement-learning, with the abstract giving enough grounded detail for follow-up reading and comparison.

English Summary

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill^2-Bench, a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels....

原文摘要 / Source Excerpt

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Abstract

Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different reasoning skills and depend on earlier outputs. Existing benchmarks often evaluate individual skills, lacking a principled way to measure how well a model switches between skills. We address this gap from both the evaluation and training sides. We introduce Skill Entropy, a measure of the difficulty of switching from one skill to another. We then propose Skill^2-Bench, a benchmark of cross-skill long-horizon tasks built over 558 skills across 9 verifiable and open-ended domains. Each task is assigned a task-level skill-entropy score and grouped into three difficulty levels. Evaluating 8 frontier and 4 open-source models on Skill^2-Bench reveals a skill-switching gap: accuracy decreases on higher-entropy tasks. We then turn skill entropy from a benchmark scale into a training signal. We propose Skill-Entropy RL, an RL framework where the model predicts not only the answer at each step but also the skill used to produce it. The reward combines step-level correctness with a skill-entropy reward that measures the alignment between the model-predicted skill sequence and the gold skill sequence. On Qwen3-4B-Instruct and Qwen3-1.7B, Skill-Entropy RL improves the Skill^2-Bench score from 34.4% to 68.4% and from 14.6% to 40.1%, respectively, outperforming competitive baselines. The same pipeline can be applied to off-the-shelf training data such as OpenR1-Math, indicating that skill entropy is a reusable training signal. Code available at: https://github.com/Gen-Verse/Skill-Entropy-RL