Agent 与自动化 4.0 · 优秀 2026-08-11 · 论文

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

论文提出 GUI Visual Grounding 的 test-time 自演化闭环:Exploration Evaluation Reflection InternalizationMLLM Reflector 给预测打分并生成 reasoning reflection,Reflection-Guided On-Policy Self-Distillation 把高层反思转成 token-level 监督,Contrastive Calibration 防止失败探索污染前缀,六个 benchmark 平均提升 7.4%

打开原文回到归档

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

Source: https://arxiv.org/abs/2608.11191
Author: Shiyu Xuan, Zechao Li
Original date: 2026-08-11
Added by: AAIF daily-intake-evening 2026-08-13

摘要

论文提出 GUI Visual Grounding 的 test-time 自演化闭环:Exploration → Evaluation → Reflection → Internalization。MLLM Reflector 给预测打分并生成 reasoning reflection,Reflection-Guided On-Policy Self-Distillation 把高层反思转成 token-level 监督,Contrastive Calibration 防止失败探索污染前缀,六个 benchmark 平均提升 7.4%。

English Summary

GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration. To overcome this, we propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. Specifically, the agent first explores unseen interfaces by predicting grounding coordinates for given instructions. To evaluate these explorations, we introduce an MLLM-based Reflector to assess the generated results and provide the corresponding reasoning reflections. To internalize reflection knowledge into the model weights, we propose Reflection-Guided On-Policy Self-Distillation, which translates high-level reasoning into dense token-level supervision via a conditioned self-teacher. Furthermore, we design a Contrastive Calibration method to prevent incorrect auto-regressive prefixes from corrupting the supervisory signals during failed explorations. Extensive experiments across six benchmarks demonstrate our framework's effectiveness, achieving an average accuracy improvement of 7.4% over the base model. To the best of our knowledge, this is the first work to successfully exploit on-policy self-distillation for test-time adaptation in GUI visual grounding. By filling the gap in post-deployment adaptation, our framework completes the self-evolving capability of GUI agents. The code will be released.

入库理由

  • quality_score: 4
  • category: agents
  • tags: gui-grounding, mobile-agent, self-evolving, visual-grounding
  • one_liner: GUI grounding 可以在部署后通过反思、自蒸馏和校准形成自演化闭环。

Obsidian evidence excerpt

md`
- Evidence: `/Users/gracker/.hermes/evidence/paper-pipeline/2026-08-13/`

## 今日论文速报

今天 arXiv recent 覆盖 `cs.AI / cs.CL / cs.LG / cs.CV / cs.RO / cs.MA` 六个分类共 254 篇去重提交。本轮最有信号的方向是 **Agent 记忆与自演化基础设施**——多条线从不同角度在做同一件事:把"retrieve-only"换成"compile + skill + provenance"。`SkillZip / Muscle Memory / GeoForge / MAP-Graph / EvoMem` 形成今天的"memory 范式切换"主线;`Self-Evolving GUI Grounding` + `SPIEval` 把这条主线拉到 GUI/移动端;`ReRound / Gated VLA-Cache` 提供 on-device 推理的量化与缓存路径。安全侧一条新的攻击面:`Trajectory Backdoor Attack` 直接攻击 self-evolving skill 的可信演化管道。

1. **SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure** — arXiv:2608.11079。把 self-evolving agent 累积的 skill 当成"typed contract"——名字/描述/工作流/工具契约/输出字段/例外规则——用 typed minimum description-length 目标一次性压缩:重复规则 state once at scope,重复动作序列 factor into shared procedure,例外保留为 explicit exceptions。Zip-on-Write 模式支持 incremental 演化不重放任务。压缩率高、保留 unique rare rules by construction。
   来源:https://arxiv.org/abs/2608.11079

2. **Muscle Memory for Agents: Compile not Merely Retrieve** — arXiv:2608.08995。主张把"反复出现的用户意图"编译成 purpose-built specialist agent,而不是 retrieve-then-orchestrate。Harvest→Analyze→Augment→Evaluate 四阶段管线,从对话历史中分别挖出 behavioral pattern 和 task pattern,发出的 specialist 配 two-stage trigger matching。90 held-out scenario 上 specialist 触发时 88.9% 胜率,+2.05 personalisation gain,accuracy 损失仅 -0.28(1-4 scale)。
   来源:https://arxiv.org/abs/2608.08995

3. **GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning** — arXiv:2608.10494。Training-free self-evolving 框架,把完成的轨迹压成结构化 nonparametric execution state——Workflow Graph Memory(全局操作顺序)+ Action-Level Experiences(局部纠错)+ Adapted Skill SOP(程序与数据约束)。执行、蒸馏、复用三段循环里 backbone LLM 不变。多个 geospatial benchmark 上同时拉高 task accuracy 和 tool-use trajectory quality。
   来源:https://arxiv.org/abs/2608.10494

4. **MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows** — arXiv:2608.10509。把 agent、source、memory、claim、action 全部建模成 typed execution graph:lineage tracing + permission-ineligible record exclusion + semantic similarity × multiplicative path trust reranking + risk-sensitive action gate。2,700 合成任务

arXiv metadata / abstract

GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration. To overcome this, we propose a Test-Time Self-Evolving framework that enables models to improve after deployment without human-annotated ground truth. It constructs a closed-loop of Exploration, Evaluation, Reflection, and Internalization. Specifically, the agent first explores unseen interfaces by predicting grounding coordinates for given instructions. To evaluate these explorations, we introduce an MLLM-based Reflector to assess the generated results and provide the corresponding reasoning reflections. To internalize reflection knowledge into the model weights, we propose Reflection-Guided On-Policy Self-Distillation, which translates high-level reasoning into dense token-level supervision via a conditioned self-teacher. Furthermore, we design a Contrastive Calibration method to prevent incorrect auto-regressive prefixes from corrupting the supervisory signals during failed explorations. Extensive experiments across six benchmarks demonstrate our framework's effectiveness, achieving an average accuracy improvement of 7.4% over the base model. To the best of our knowledge, this is the first work to successfully exploit on-policy self-distillation for test-time adaptation in GUI visual grounding. By filling the gap in post-deployment adaptation, our framework completes the self-evolving capability of GUI agents. The code will be released.