ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- ID: 80ebb758
- Original URL: https://arxiv.org/abs/2608.05102
- PDF: https://arxiv.org/pdf/2608.05102v1
- Author(s): Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen
- Date: 2026-08-05
- Category: cs.AI
- Source type: paper
- Tags: search-agents, credit-assignment, reinforcement-learning, browsecomp, long-horizon
- Quality score: 4/5
- Fetched at: 2026-08-07T04:20:20+00:00
- Obsidian evidence: OpenCLI arXiv metadata backfill
中文导读
ABSeeker 聚焦长程搜索 agent 的训练信号问题:现有 SFT/RL 往往把整条轨迹中的步骤统一处理,无法区分有用错误或冗余行动论文提出 Answer-Backtracked Credit Assignment,先从正确答案回溯中间线索,再对每个搜索步骤做 clue-anchored scoring基于 Qwen3.5-4B 和 8.5k 样本,摘要报告 BrowseComp 37.3%BrowseComp-ZH 39.1%,加入上下文管理后达 55.3% 和 52.9%
为什么值得关注
This fills a high-score (4/5) AAIF content gap around search-agents, credit-assignment, reinforcement-learning, browsecomp, with the abstract giving enough grounded detail for follow-up reading and comparison.
English Summary
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions....
原文摘要 / Source Excerpt
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
- arXiv: https://arxiv.org/abs/2608.05102
- PDF: https://arxiv.org/pdf/2608.05102v1
- Authors: Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen
- Published: 2026-08-05
- Updated: 2026-08-05
- Categories: cs.AI
Abstract
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions. Specifically, given a potentially obscure query and its corresponding ground-truth answer, ABC first performs Answer-Backtracked Clue Recovery, which traces back from the answer to recover intermediate clues required to solve the question. It then applies Clue-Anchored Step Scoring to evaluate each search step against these clues, converting sparse binary outcome supervision into dense step-level rewards. Based on these rewards, we develop ABC-SFT, which reweights the loss of each turn, and ABC-GRPO, which uses the step-level scores as rewards in GRPO. Building on this framework, we train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH. With context management, the scores further improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (approximately 30B). These results demonstrate the effectiveness of answer-backtracked step-level credit assignment for training long-horizon search agents.