Agent 与自动化 4.0 · 优秀 2026-08-20 · 论文

Inducing Task Models from Computer-Use Traces

提出 Task Model Induction (TMI):从被动录制的屏幕截图与键鼠轨迹中自动发现潜在任务解缠并发活动,并为每个任务归纳层次化目标模型 + 控制流过程模型在人类与 agent 轨迹上,TMI 以 0.974 的一致率恢复交错任务重建 74.9% 的执行步骤,远超最强工作流归纳基线;由任务模型导出的技能让留出任务准确率再提升 30.0%对进入真实工作流的 computer-use agent 而言,这是把日常操作痕迹变成可审计可复用技能资产的系统化路径

打开原文回到归档

Inducing Task Models from Computer-Use Traces

  • arXiv: 2608.20319
  • Authors: Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen, Diyi Yang
  • Published: 2026-08-20; categories: cs.CL, cs.AI

Abstract

Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and reusable models of how everyday work is done. Such models matter as computer-use agents enter real work, where agents need to learn how tasks are actually performed, and organizations need to audit and reuse that knowledge. However, inducing such task models is challenging, as activity is observed only as low-level events and real-world work is multi-threaded with interleaved goals. Existing methods assume a given task or a single workflow, and produce step-level summaries rather than structured task models. We introduce Task Model Induction (TMI), which (i) discovers the latent tasks in an unconstrained trace, disentangling concurrent activity, and (ii) for each latent task, induces a task model pairing a hierarchical objective model of recursive goal decomposition with a procedure model of the control flow that organized the execution. Intrinsically, on controlled human and agent trajectories, TMI recovers interleaved tasks with 0.974 agreement against ground-truth groupings and reconstructs 74.9% of the observed execution steps, far more than the strongest workflow induction baseline. Extrinsically, skills derived from TMI's task models improve held-out task accuracy by 30.0% over the strongest baseline.

Why it matters (AAIF scan)

提出 Task Model Induction (TMI):从被动录制的屏幕截图与键鼠轨迹中自动发现潜在任务解缠并发活动,并为每个任务归纳层次化目标模型 + 控制流过程模型在人类与 agent 轨迹上,TMI 以 0.974 的一致率恢复交错任务重建 74.9% 的执行步骤,远超最强工作流归纳基线;由任务模型导出的技能让留出任务准确率再提升 30.0%对进入真实工作流的 computer-use agent 而言,这是把日常操作痕迹变成可审计可复用技能资产的系统化路径

Source: https://arxiv.org/abs/2608.20319
Captured: 2026-08-22 (AAIF content-fetcher)