ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize
Source: https://arxiv.org/abs/2609.04197
Authors: Lihao Liu, Peng Tang, Kunwar Yashraj Singh, Shabnam Ghadar
Published: 2026-09-03
Categories: cs.CL, cs.AI
PDF: https://arxiv.org/pdf/2609.04197
Abstract
Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3x longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; Propose generates candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection. On seven public NLP benchmarks - Tweet, MMLU, GSM8K, HotpotQA, ScoNe, HoVer, and PUPA - ESPO improves average accuracy by +3.76 pp over the state-of-the-art (74.67% vs 70.91% for GEPA), matching or exceeding GEPA on every dataset while producing prompts 47% shorter (1,004 vs 1,878 chars) and faster at inference. Cross-model experiments across four additional student models (Gemma 3 12B, Mistral 14B, Qwen3 32B, Claude Haiku 4.5) show ESPO yields the best average accuracy on every model tested, with the largest gap on Qwen3 GSM8K (15.00% to 91.40%). A generalization bound (Appendix) grounds each phase in a corresponding term of the test-time gap, and the ablation confirms a key prediction: adding diversity without bootstrap selection actually hurts performance (-1.20%).
Key Points
- Evolutionary prompt optimizers such as GEPA suffer prompt bloat: each iteration appends rules and caveats, producing prompts up to 3x longer yet no more accurate.
- ESPO traces the bloat to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection.
- Three phases: Diagnose clusters all training errors into structural patterns in one round; Propose generates candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection.
- On seven public NLP benchmarks (Tweet, MMLU, GSM8K, HotpotQA, ScoNe, HoVer, PUPA), ESPO improves average accuracy by +3.76 pp over the state of the art (74.67% vs 70.91% for GEPA), matching or exceeding GEPA on every dataset.
- Prompts come out 47% shorter (1,004 vs 1,878 chars) and faster at inference; cross-model experiments on Gemma 3 12B, Mistral 14B, Qwen3 32B, and Claude Haiku 4.5 show the best average accuracy on every model tested, with the largest gap on Qwen3 GSM8K (15.00% to 91.40%).
- A generalization bound grounds each phase in a corresponding term of the test-time gap, and the ablation confirms a key prediction: adding diversity without bootstrap selection actually hurts performance (-1.20%).
中文概要
针对 GEPA 类进化式 prompt 优化器的 prompt 膨胀问题(每轮追加规则与告警,prompt 最长 3 倍而准确率无提升),作者归因于三个缺陷:错误观察不完整搜索多样性受限选择不可靠ESPO 把优化拆成 Diagnose(一轮内把全部训练错误聚类成结构化模式)Propose(四种互补策略独立偏置生成候选)Select(bootstrap 稳定的选择)三阶段EMNLP 2026 收录,2026-09-03 提交 cs.CL,实验细节以论文正文为准