Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
- ID: dbde5cbb
- 原文链接: https://arxiv.org/abs/2608.24842
- PDF: https://arxiv.org/pdf/2608.24842v1
- 作者: Miao Liu, Zhizhe Liu
- 日期: 2026-08-25
- 更新: 2026-08-25
- 分类: learning
- 来源类型: paper
- 标签: llm, retrieval, long-context, financial-analysis, evaluation, arxiv
- 质量评分: 4/5
- 抓取时间: 2026-08-27T05:21:12Z
中文导读
LLM 正在被当作 AI 分析师处理金融信息披露并辅助投资决策,但这类系统通常按能检索到什么评估,而不看检索到的信息是否真正改变了判断论文识别出长上下文金融分析中的检索-整合鸿沟:固定目标公司信息不变只把无关上下文从 2,000 提高到 128,000 token 时,风险披露对投资判断的影响会跌到实验噪声底,而直接检索仍然准确该模式在多个模型族与判断任务上复现,也在从真实 10-K 文件中删去真实披露的反事实实验中出现;更强的模型只会推迟而不会消除这一鸿沟因果记忆干预显示,压缩摘要和源文查找二者结合才能把披露信息传递进决策
为什么值得关注
读到了不等于用上了:长上下文金融分析存在检索-整合鸿沟,无关上下文涨到 128K token 后风险披露对投资判断的影响只剩噪声
关键信息
- 论文标题: Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
- 作者: Miao Liu, Zhizhe Liu
- arXiv: https://arxiv.org/abs/2608.24842
- 发布时间: 2026-08-25
- arXiv 分类: cs.CL, cs.AI
- 备注: 无
- 关联标签: llm, retrieval, long-context, financial-analysis, evaluation, arxiv
- Context length sweep 2k -> 128k unrelated tokens; risk disclosure influence falls to noise floor while retrieval stays accurate
- Chunk-and-summarize pipelines evict relevant info; targeted structured restatement adjacent to the decision restores influence
English Abstract
Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.
Obsidian Notes
- Metadata and abstract fetched via
opencli arxiv paper 2608.24842 -f json(2026-08-27T05:21:12Z); response parsed list-or-dict tolerant. - 中文导读与价值判断锚定在条目已有摘要与论文摘要上,未补充摘要之外的实验细节。
- Canonical backfill page for existing entry
dbde5cbbindata/entries.json; generated by AAIF content-fetcher (2026-08-27), no push.