EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
- ID: 0176a3ac
- 原文链接: https://arxiv.org/abs/2609.40340
- PDF: https://arxiv.org/pdf/2609.40340v1
- 作者: Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang
- 日期: 2026-09-30
- 更新: 2026-09-30
- 分类: agents
- 来源类型: paper
- 标签: evolutionary-search, web-search, retrieval, scientific-discovery
- 质量评分: 4/5
- 抓取时间: 2026-10-02T12:30:30.134235+00:00
中文导读
LLM 进化式搜索在进展依赖模型未知的外部知识时会停滞,直接挂上 web 搜索工具又会反复抓回同样的页面EvoDuet 用双层优化让解与搜索查询共同进化模型参数冻结:每轮由 retrieval gate 评估知识缺口并决定取新文档复用已存文档还是不检索;内循环细化查询并按预测解分数给文档排序,外循环从文档并行生成候选并记录评估结果供后续搜索使用21 个任务每轮单候选设定下,OpenEvolve 归一化发现增益从 74.1% 提到 78.0%(GPT-5.6-Luna)61.3% 提到 82.3%(Gemini-3.8-Flash),Qwen3.5-9B 无收益;最佳运行在 8 个任务上超过此前报告最高分,并兼容 Top-KEvoX 等其他 scaffold
为什么值得关注
让搜索查询和解一起进化:retrieval gate + 双层优化把 OpenEvolve 发现增益最高提 21 个点
关键信息
- 论文标题: EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
- 作者: Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang
- arXiv: https://arxiv.org/abs/2609.40340
- arXiv 分类: cs.CL
- 发布时间: 2026-09-30
- 关联标签: evolutionary-search, web-search, retrieval, scientific-discovery
English Abstract
Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.
English Summary
EvoDuet is a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters: a retrieval gate decides whether to fetch new documents, reuse stored ones, or proceed without; an inner loop refines queries and ranks documents by predicted solution scores; an outer loop generates candidates in parallel and records evaluated outcomes for later searches. Across 21 tasks with one candidate per iteration, OpenEvolve's normalized discovery gain rises from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash (no benefit for Qwen3.5-9B), surpassing prior best scores on eight tasks and transferring to other scaffolds such as Top-K and EvoX.
Obsidian Notes
- Content grounded in an opencli fetch during content-fetcher backfill (papers via
opencli arxiv paper -f json; articles viaopencli web read). - 中文导读 block is the entry's summary_zh (verbatim); the value note is the entry's one_liner; English blocks are verbatim excerpts from the fetched source. No claims beyond the fetched metadata/source text were added.