模型与实验室 5.0 · 必读 2026-08-13 · 论文

Intern-S2-Preview: Scientific Agentic Foundation Model

Intern-S2-Preview 是面向科学发现场景的智能体基座模型系列:先在渲染科学文档图文交错语料与多类科学语料上做多模态预训练,再用统一后训练管线(SFT可扩展多任务 RL黑盒/白盒 agentic RL在线蒸馏)训练工程侧贡献密集:partial rollout 加离策略校正自适应长度正则在线投机解码鲁棒多任务优化与 trace 级经验组装397B 版本将时序建模从长序列理解扩展到数值预报,并以 Memory Decoder 作为冻结主干外的记忆增强路径做快速科学专化科学多模态智能体与通用基准上取得有竞争力或领先的成绩收录理由:当前公开的最完整科学 agentic 基座模型训练配方,agentic RL 工程细节密度高

打开原文回到归档

Intern-S2-Preview: Scientific Agentic Foundation Model

  • ID: 18a7a4f7
  • 原文链接: https://arxiv.org/abs/2608.13505
  • PDF: https://arxiv.org/pdf/2608.13505v1
  • 作者: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, 等 125 位作者(完整名单见原文)
  • 日期: 2026-08-13
  • 更新: N/A
  • 分类: models
  • 来源类型: paper
  • 标签: scientific-ai, agentic-rl, multimodal, foundation-model, arxiv
  • 质量评分: 5/5
  • 抓取时间: 2026-08-17T04:23:51+00:00Z

中文导读

Intern-S2-Preview 是面向科学发现场景的智能体基座模型系列:先在渲染科学文档图文交错语料与多类科学语料上做多模态预训练,再用统一后训练管线(SFT可扩展多任务 RL黑盒/白盒 agentic RL在线蒸馏)训练工程侧贡献密集:partial rollout 加离策略校正自适应长度正则在线投机解码鲁棒多任务优化与 trace 级经验组装397B 版本将时序建模从长序列理解扩展到数值预报,并以 Memory Decoder 作为冻结主干外的记忆增强路径做快速科学专化科学多模态智能体与通用基准上取得有竞争力或领先的成绩

为什么值得关注

Intern-S2-Preview 是面向科学发现场景的智能体基座模型系列:先在渲染科学文档图文交错语料与多类科学语料上做多模态预训练。

收录理由:当前公开的最完整科学 agentic 基座模型训练配方,agentic RL 工程细节密度高

关键信息

  • 论文标题:Intern-S2-Preview: Scientific Agentic Foundation Model
  • 作者:Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, 等 125 位作者(完整名单见原文)
  • arXiv:https://arxiv.org/abs/2608.13505
  • 发布时间:2026-08-13
  • arXiv 分类:cs.LG, cs.CL, cs.CV
  • 关联标签:scientific-ai, agentic-rl, multimodal, foundation-model, arxiv

English Abstract

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.

English Summary

Intern-S2-Preview is a series of scientific agentic foundation models for multimodal scientific understanding, reasoning, generation, and long-horizon tasks. Training combines scientific multimodal pre-training with a unified post-training pipeline: SFT, scalable multi-task RL, black- and white-box agentic RL, and on-policy distillation, supported by partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, and trace-aware experience assembly. The 397B variant extends time-series modelling to numerical forecasting, with a Memory Decoder as a separate memory-augmented path for rapid scientific specialization over the frozen backbone; evaluations show competitive or leading results across scientific, multimodal, agentic, and general benchmarks.

Obsidian Notes

  • 内容由 opencli arxiv paper 拉取 arXiv 元数据与摘要生成。
  • 中文导读与价值判断锚定于条目与论文摘要,未补充摘要之外的内容。