AI 编程 4.0 · 优秀 2026-08-05 · 论文

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

OctoLong 针对长上下文模型在 agentic workflows 和代码仓库理解中缺少远距离依赖语料的问题,用 AST parserlanguage server backend 和 package manager 构建跨仓库代码引用检索管线,生成百万 token 级的依赖密集上下文摘要报告,在约 50B token 中训练并仅用 12% OctoLong 数据替换传统语料,就提升长距离检索长期状态跟踪repo-level 代码理解和下游 agent 任务

打开原文回到归档

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

  • ID: 18262696
  • Original URL: https://arxiv.org/abs/2608.05141
  • PDF: https://arxiv.org/pdf/2608.05141v1
  • Author(s): Indraneil Paul, Falko Helm, Goran Glavaš, Iryna Gurevych
  • Date: 2026-08-05
  • Category: cs.AI, cs.LG, cs.SE
  • Source type: paper
  • Tags: long-context, code-agents, context-engineering, repository-understanding, training-data
  • Quality score: 4/5
  • Fetched at: 2026-08-07T04:20:20+00:00
  • Obsidian evidence: OpenCLI arXiv metadata backfill

中文导读

OctoLong 针对长上下文模型在 agentic workflows 和代码仓库理解中缺少远距离依赖语料的问题,用 AST parserlanguage server backend 和 package manager 构建跨仓库代码引用检索管线,生成百万 token 级的依赖密集上下文摘要报告,在约 50B token 中训练并仅用 12% OctoLong 数据替换传统语料,就提升长距离检索长期状态跟踪repo-level 代码理解和下游 agent 任务

为什么值得关注

This fills a high-score (4/5) AAIF content gap around long-context, code-agents, context-engineering, repository-understanding, with the abstract giving enough grounded detail for follow-up reading and comparison.

English Summary

Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies. In this work, we introduce OctoLong, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of dependency-rich code contexts of millions of tokens in length....

原文摘要 / Source Excerpt

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

Abstract

Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies. In this work, we introduce OctoLong, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of dependency-rich code contexts of millions of tokens in length. We then train OctoLong-Instruct, a suite of capable long-context open LMs, derived from base models ranging in size from 600M to 14B parameters, via context-extension mid-training on a ~50B-token mixture containing ~6.2B tokens of OctoLong code contexts, followed by ~10B tokens of instruction tuning. Our training ablations and experimental evaluations against 18 state-of-the-art open-weight long-context LMs show that supplanting just 12% of traditional context-extension corpora with OctoLong data yields substantial gains in long-range retrieval, long-term state tracking, repository-level code understanding, and downstream agentic tasks, while also enhancing API usage in short-context coding scenarios.