Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Source: https://arxiv.org/abs/2608.06305
PDF: https://arxiv.org/pdf/2608.06305v1
Content fetched: 2026-08-09T15:34:20.317851+00:00
Grounding: OpenCLI arXiv metadata and abstract; Obsidian evidence: OpenClaw定时任务/论文流水线/2026-08-09-论文流水线.md
Metadata
- Author(s): Sagar Tamang, Ayush Vyas, Tabarakul Hazarika
- Original date: 2026-08-06
- Platform: arxiv
- AAIF quality score: 5
中文摘要
论文批评长文档 RAG 默认 top-k embedding 在表格密集文档上结构性失效:政府财报中大量表格行、相似数字和跨行单位会被 chunk 边界切开。READ 暴露规范化词法搜索、结构导航、有限跨度读取三种确定性操作,通过 MCP 给 agent 使用,让检索轨迹可回放。51 个验证问题上 READ 答对 58.8%,dense retrieval 为 15.7%,但作者也明确 BM25 与 READ 统计上不可区分。
English Summary
READ replaces opaque top-k embedding retrieval with deterministic operations: normalized lexical search, structural navigation, and bounded span reads exposed over MCP. On a 780-page financial report and 51 verified questions, READ outperforms dense retrieval while producing replayable audit trails, though BM25 is statistically indistinguishable from READ.
Abstract excerpt
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is structurally unsound, and we make the argument measurable. We propose READ (Reliable Embedding-free Agentic Document-search), in which an agent reads the raw document through three deterministic operations -- normalized lexical search, structural navigation, and bounded span reads -- exposed over the Model Context Protocol, so a trajectory is a replayable audit trail, not an opaque similarity score.
Why it matters for AAIF
表格密集长文档的 RAG 需要可回放操作轨迹,不能只相信 top-k embedding 相似度。