Agent 与自动化 4.0 · 优秀 2026-08-23 · X

Jerry Liu: two-pass document processing is the default for agent harnesses

LlamaIndex CEO Jerry Liu 总结当前 agent harness(CodexCowork)的 RAG 新趋势:data room 级知识任务默认两遍文档处理第一遍廉价 OSS 解析扫全部文件支撑检索,第二遍只对命中页做 JIT VLM 精读开箱即用三坑:通用大模型当 OCR 贵且 grounding 弱轻量一程对复杂版式不够agent 手写图表/bbox 代码抬账单;对应工具面是 LiteParse(OSS 一程)与按页调用的 LlamaParse(二程)

打开原文回到归档

Jerry Liu: two-pass document processing is the default for agent harnesses

Source: https://x.com/jerryjliu0/status/2091564183922077885
Author: Jerry Liu (@jerryjliu0), LlamaIndex CEO — 19K views, 220 likes at fetch time

(5) Jerry Liu on X: "The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents:

1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across https://t.co/5BHqmKmRZy" / X

发布时间: 2026-08-23T16:32:15.000Z
原文链接: https://x.com/jerryjliu0/status/2091564183922077885

[

](https://x.com/jerryjliu0)

[

Jerry Liu

](https://x.com/jerryjliu0)

[

@jerryjliu0

](https://x.com/jerryjliu0)

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: \* Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding \* The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. \* The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within

@llama\_index

to help any agent do two-pass document processing with higher accuracy and lower cost. We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: https://github.com/run-llama/liteparse… LlamaParse: https://cloud.llamaindex.ai All the relevant docs, including MCP, are here: https://developers.llamaindex.ai/llamaparse/for\-agents/mcp/…

<video src="blob:https://x.com/a87bc0d2-40f3-4262-84ff-1d8c5f1f2366&quot; controls poster="https://pbs.twimg.com/amplify_video_thumb/2091564028007157760/img/Ib0Yc0RUiUmeeLWm.jpg&quot;&gt;&lt;/video&gt;

12:32 AM · Aug 24, 2026

[

19K

Views](https://x.com/jerryjliu0/status/2091564183922077885/analytics)

32

31

220

280

Relevant

View quotes

证据摘录(Obsidian 调研 2026-08-24)

面向 data room / 文件夹级知识任务,默认架构应是:Pass1 廉价本地解析 + 检索定位页码 → Pass2 只对命中页做 VLM/重解析。全量 VLM OCR 只适合页数少、几乎每页都要读图的短任务。

配套工具面(官方文档可核对):LiteParse(本地 OSS,Rust,complexity score 先评分再路由);LlamaParse MCP 的 estimateFileComplexity / parseWithLiteParse(不计 Parse credits)/ parseFile 三工具路由。