TREK:面向 LLM 智能体复杂旅行规划的推理与评测套件
Source: arXiv:2607.26977 · Category: cs.CL · Authors: Jinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu, Yanyu Chen, Yaoman Li, Irwin King
TL;DR
旅行规划是工具型 LLM 智能体的严苛压力测试:一份可用的行程必须同时在多个纬度上同时正确,每个航班、酒店与景点都要存在且可预订,天数路线要物理可走,总价要在预算内,还要能满足默默提出、未明说的个人需求。现有基准一次只考一个纬度、用软标准或 LLM 裁判评分,无法证明返回的行程可执行,也不可复现、不可审计。TREK 要求一份行程同时约束正确、无幻觉、时空可执行、预算合规、响应未明人设需求。基准含 800 个多约束任务:533 可行、 267 可证不可行,均附详细的类型化原因(路线 / 实体 / 预算);以一个 212,530 条记录、覆盖 375 个城市、 13 个人设的一致性知识库为背景,提供生产级的 RESTful API 工具沙盒。评分器是全定制、规则驱动的,不含 LLM 裁判;每个任务都附一份人工验证、能在同一评分器下得到满分的金标准。评估 15 个智能体以 9 个约束维度:最强的 GPT-5.6 仅在 46.2% 的可解任务上产出全部可行的行程,中位数 6.6%,最低 0.0%;“满足未明说的个人需求”是全局瓶颈,连前沿模型亦未解决。数据、工具沙盒、评分器与智能体代码全部开源可复现。
为什么重要
以旅行规划为代理的多约束任务:“多个纬度同时正确”是智能体评测现状的谴痛。TREK 提供全规则、可重放、人工验证金标准的评分器,最强智能体仅 46.2%、中位 6.6%;“未明人设需求”是全场瓶颈。
原文摘要(English)
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, the total must clear a budget, and the plan must serve a traveler whose needs are only partly stated. Existing agent benchmarks reward these properties one at a time and grade the final output with soft or LLM-judged rubrics, which cannot certify that a returned plan is executable and are neither reproducible nor auditable. We introduce TREK (Travel Reasoning and Evaluation Kit), a benchmark for feasible itinerary synthesis: producing a single plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to the traveler's unstated persona needs. TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas, served through a production-style tool sandbox of validated RESTful APIs. Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator, so the ceiling is demonstrably achievable and every remaining gap is an agent limitation rather than scorer strictness. Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks, with a median of 6.6% and a floor of 0.0%; satisfying travelers' unstated needs emerges as the universal bottleneck, unsolved even at the frontier. We release the dataset, tool sandbox, deterministic evaluator, and agent code as a fully reproducible benchmark.
基本信息
| 项 | 值 | |------|------| | 论文 ID | 2607.26977 | | 主分类 | cs.CL | | 发表 | 2026-07-29 | | 更新 | 2026-08-10 | | 作者 | Jinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu, Yanyu Chen, Yaoman Li, Irwin King | | 备注 | Code, data, and evaluator: https://github.com/TonyQJH/TREK-A-Travel-Reasoning-and-Evaluation-Kit-for-LLM-Agents-in-Complex-Trip-Planning | | PDF | 2607.26977 | | abs | https://arxiv.org/abs/2607.26977 |
参考
- arXiv abs: <https://arxiv.org/abs/2607.26977>
- arXiv PDF: <https://arxiv.org/pdf/2607.26977v2>