SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Source: https://arxiv.org/abs/2608.10692
Author: Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
Original date: 2026-08-11
Added by: AAIF daily-intake-evening 2026-08-13
摘要
SPIEval 给手机助手 LLM 做了结构化评测:250 个任务、4,335 条个人记录、10 个 app、21 个工具,覆盖 reasoning、disambiguation、integration、preference inference 与 multi-intent decomposition。GPT-5.5 最高 57.3%,失败的 79% 来自 inaccurate information localization,说明移动端 agent 的核心难点仍是持续检索与定位。
English Summary
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.
入库理由
- quality_score: 5
- category: agents
- tags: mobile-assistant, agent-evaluation, personal-information, tool-use
- one_liner: SPIEval 把手机助手失败定位到 scattered personal information 的检索与定位能力。
Obsidian evidence excerpt
.hermes/evidence/paper-pipeline/2026-08-13/`
## 今日论文速报
今天 arXiv recent 覆盖 `cs.AI / cs.CL / cs.LG / cs.CV / cs.RO / cs.MA` 六个分类共 254 篇去重提交。本轮最有信号的方向是 **Agent 记忆与自演化基础设施**——多条线从不同角度在做同一件事:把"retrieve-only"换成"compile + skill + provenance"。`SkillZip / Muscle Memory / GeoForge / MAP-Graph / EvoMem` 形成今天的"memory 范式切换"主线;`Self-Evolving GUI Grounding` + `SPIEval` 把这条主线拉到 GUI/移动端;`ReRound / Gated VLA-Cache` 提供 on-device 推理的量化与缓存路径。安全侧一条新的攻击面:`Trajectory Backdoor Attack` 直接攻击 self-evolving skill 的可信演化管道。
1. **SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure** — arXiv:2608.11079。把 self-evolving agent 累积的 skill 当成"typed contract"——名字/描述/工作流/工具契约/输出字段/例外规则——用 typed minimum description-length 目标一次性压缩:重复规则 state once at scope,重复动作序列 factor into shared procedure,例外保留为 explicit exceptions。Zip-on-Write 模式支持 incremental 演化不重放任务。压缩率高、保留 unique rare rules by construction。
来源:https://arxiv.org/abs/2608.11079
2. **Muscle Memory for Agents: Compile not Merely Retrieve** — arXiv:2608.08995。主张把"反复出现的用户意图"编译成 purpose-built specialist agent,而不是 retrieve-then-orchestrate。Harvest→Analyze→Augment→Evaluate 四阶段管线,从对话历史中分别挖出 behavioral pattern 和 task pattern,发出的 specialist 配 two-stage trigger matching。90 held-out scenario 上 specialist 触发时 88.9% 胜率,+2.05 personalisation gain,accuracy 损失仅 -0.28(1-4 scale)。
来源:https://arxiv.org/abs/2608.08995
3. **GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning** — arXiv:2608.10494。Training-free self-evolving 框架,把完成的轨迹压成结构化 nonparametric execution state——Workflow Graph Memory(全局操作顺序)+ Action-Level Experiences(局部纠错)+ Adapted Skill SOP(程序与数据约束)。执行、蒸馏、复用三段循环里 backbone LLM 不变。多个 geospatial benchmark 上同时拉高 task accuracy 和 tool-use trajectory quality。
来源:https://arxiv.org/abs/2608.10494
4. **MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows** — arXiv:2608.10509。把 agent、source、memory、claim、action 全部建模成 typed execution graph:lineage tracing + permission-ineligible record exclusion + semantic similarity × multiplicative path trust reranking + risk-sensitive action gate。2,700 合成任务上 94.96% task success / 72.70% e
arXiv metadata / abstract
- arXiv id: 2608.10692
- authors: Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
- published: 2026-08-11
- updated: 2026-08-11
- categories: cs.CL, cs.AI
- PDF: https://arxiv.org/pdf/2608.10692v1
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.