模型与实验室 4.0 · 优秀 2026-07-28 · 文章

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

提出一个包含14个能力域和91个子技能的多层分类体系,覆盖原始层构建层和整合层,以人类认知科学而非LLM架构为指导作者筛选了ACLAAAIICML和NeurIPS在2023-2025年间发表的31505篇论文,映射了其中15934篇LLM论文研究注意力集中在语言语义能力(22.3%)推理(21.3%)规划决策(13.5%)和感知(12.3%),有六个能力域不到2%心智理论与社会推理的共现lift高达30.84

打开原文回到归档

From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models

  • ID: ffc376ee
  • 原文链接: https://arxiv.org/abs/2607.22182
  • PDF: https://arxiv.org/pdf/2607.22182v1
  • 作者: Shixin Fang, Jiachen Wo, Wenjuan Qin, Sihang Jiang, Yanghua Xiao
  • 日期: 2026-07-24
  • 更新: 2026-07-24
  • 分类: models
  • 来源类型: article
  • 标签: taxonomy, capabilities, evaluation, benchmark-mapping, cognitive-science
  • 质量评分: 4/5
  • 抓取时间: 2026-07-29T12:23:37Z
  • arXiv categories: cs.CL, cs.AI
  • comment: 34 pages, 5 figures, 20 tables

中文导读

提出一个包含14个能力域和91个子技能的多层分类体系,覆盖原始层构建层和整合层,以人类认知科学而非LLM架构为指导作者筛选了ACLAAAIICML和NeurIPS在2023-2025年间发表的31505篇论文,映射了其中15934篇LLM论文研究注意力集中在语言语义能力(22.3%)推理(21.3%)规划决策(13.5%)和感知(12.3%),有六个能力域不到2%心智理论与社会推理的共现lift高达30.84

为什么值得关注

提出一个包含14个能力域和91个子技能的多层分类体系,覆盖原始层构建层和整合层,以人类认知科学而非LLM架构为指导作者筛选了ACLAAAIICML和NeurIPS在2023-2025年间发表的31505篇论文,映射了其中15934篇LLM论文研究注意力集中在语言语义能力(22.

关键信号:作者将评估单位从孤立 task/基准转向结构化 capability,并对 2023–2025 ACL/AAAI/ICML/NeurIPS 上 15,934 篇 LLM 论文做了能力映射与覆盖度审计。语义与推理占主,多个社会/心智相关域显著低表现;Theory of Mind 与 Social Reasoning 共现 lift 极高,对评估设计与训练转移假设有直接启发。

关键信息

  • 论文标题:From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models
  • 作者:Shixin Fang, Jiachen Wo, Wenjuan Qin, Sihang Jiang, Yanghua Xiao
  • arXiv:https://arxiv.org/abs/2607.22182
  • 发布时间:2026-07-24
  • arXiv 分类:cs.CL, cs.AI
  • 关联标签:taxonomy, capabilities, evaluation, benchmark-mapping, cognitive-science

English Abstract

Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify. We introduce a multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers. Human cognitive science guides capability definition and organization, not LLM architecture. Layer assignments draw on developmental precedence and hypothesized functional support, while human-origin constructs are adapted to observable model behavior. To demonstrate operational utility, we screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025 and mapped 15,934 LLM-focused papers through multi-model annotation, consensus, and arbitration. Direct research attention concentrated on Language-Semantic Competence (3,551; 22.3%), Reasoning (3,388; 21.3%), Planning and Decision-Making (2,149; 13.5%), and Perception (1,954; 12.3%), whereas six domains appeared in fewer than 2% of papers. Within domains, the most frequent subskill had a median prevalence of 97.9% and appeared in at least 90% of papers in 10 of 14 domains. Language-Semantic Competence and Reasoning formed the highest-volume pair (n = 1,864; 11.7%; lift = 2.47), whereas Theory of Mind and Social Reasoning and Interaction showed the highest lift among pairs with at least 20 co-occurrences (n = 62; lift = 30.84). By shifting the unit of analysis from isolated tasks to structured capabilities, the taxonomy supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.

English Summary

A multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers, guided by human cognitive science rather than LLM architecture. The authors screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS (2023-2025) and mapped 15,934 LLM-focused papers. Direct attention concentrates on Language-Semantic Competence (22.3%), Reasoning (21.3%), Planning (13.5%), and Perception (12.3%), while six domains appear in <2% of papers. Language-Semantic + Reasoning form the highest-volume co-occurrence pair (lift=2.47), while Theory of Mind + Social Reasoning show the highest lift among pairs with at least 20 co-occurrences (lift=30.84).

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断均锚定在条目已有摘要、论文摘要、作者、日期与分类信息上;未补充论文摘要之外的实验细节。
  • Content-fetcher backfill for active high-score entry missing canonical content file.