Agent 与自动化 4.0 · 优秀 2026-07-30 · 论文

OSReward: 跨平台计算机使用 Agent 奖励模型的标准化评测

OSReward 提出了一个现实高质量的基准测试,用于评估视觉语言模型(VLM)对计算机使用 Agent(CUA)轨迹的判定能力轨迹来自多样化 Agent 骨架执行人工验证的指令,经多阶段人工标注产生真值判决还揘导出 OSReward-Hard(难例集)和 OSReward-Multi(细粒度评分)评测发现即便是 SOTA 模型也达不到理想判定者水准,存在系统性宽容偏差

打开原文回到归档

OSReward: 跨平台计算机使用 Agent 奖励模型的标准化评测

Source: <https://arxiv.org/abs/2607.28609&gt; | 2026-07-30
Author: | Category: agents | Quality Score: 4/5
Tags: computer-use-agent, reward-model, benchmark, vlm, evaluation

English Summary

OSReward introduces a realistic, high-quality benchmark for evaluating VLM judges on computer-using agent (CUA) trajectories. Trajectories come from diverse agent backbones executing human-verified instructions, rigorously labeled through multi-stage human annotation. Derivatives include OSReward-Hard (challenge set) and OSReward-Multi (fine-grained efficiency/alignment scoring). The most comprehensive evaluation of VLM judges to date finds even SOTA models fall short of an ideal judge, sharing a systematic leniency bias.

中文概要

OSReward 提出了一个现实高质量的基准测试,用于评估视觉语言模型(VLM)对计算机使用 Agent(CUA)轨迹的判定能力轨迹来自多样化 Agent 骨架执行人工验证的指令,经多阶段人工标注产生真值判决还揘导出 OSReward-Hard(难例集)和 OSReward-Multi(细粒度评分)评测发现即便是 SOTA 模型也达不到理想判定者水准,存在系统性宽容偏差