OSReward: 跨平台计算机使用 Agent 奖励模型的标准化评测
Source: <https://arxiv.org/abs/2607.28609> | 2026-07-30
Author: | Category: agents | Quality Score: 4/5
Tags: computer-use-agent, reward-model, benchmark, vlm, evaluation
English Summary
OSReward introduces a realistic, high-quality benchmark for evaluating VLM judges on computer-using agent (CUA) trajectories. Trajectories come from diverse agent backbones executing human-verified instructions, rigorously labeled through multi-stage human annotation. Derivatives include OSReward-Hard (challenge set) and OSReward-Multi (fine-grained efficiency/alignment scoring). The most comprehensive evaluation of VLM judges to date finds even SOTA models fall short of an ideal judge, sharing a systematic leniency bias.
中文概要
OSReward 提出了一个现实高质量的基准测试,用于评估视觉语言模型(VLM)对计算机使用 Agent(CUA)轨迹的判定能力轨迹来自多样化 Agent 骨架执行人工验证的指令,经多阶段人工标注产生真值判决还揘导出 OSReward-Hard(难例集)和 OSReward-Multi(细粒度评分)评测发现即便是 SOTA 模型也达不到理想判定者水准,存在系统性宽容偏差