Intern-S2-Preview: Scientific Agentic Foundation Model
Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv, Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun, Yu Sun, Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang, Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu, Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang, Chao Zhang, Chen Zhang, Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou, Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou
Shanghai AI Laboratory
cs.LG, cs.CL, cs.CV
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 35 pages, 12 figures
Code: https://github.com/InternLM/xtuner
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Intern-S2-Preview is a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks.
Terminology
Summary
Intern-S2-Preview is a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The main model evaluated is Intern-S2-Preview-397B.
The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, a unified post-training pipeline is applied, consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks.
At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone.
Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
Architecture:
-
Memory Decoder: A separate extension model for continual domain specialization. It attaches new knowledge and specialized capabilities through external parametric memories while keeping the Intern-S2-Preview-397B backbone frozen. A lightweight token-level router determines how much the memory decoder should contribute. Training compresses retrieval-based domain evidence into a reusable parametric module using retrieval distillation and supervision from gold answer tokens. At inference, the final next-token distribution is a dynamic fusion of the backbone and memory decoder distributions.
-
Time Series Modules: The upgraded time series encoder improves long-sequence processing efficiency and multi-channel representation learning. It supports input lengths up to 300,000 time steps with 5-6× faster inference and 20% GPU memory consumption compared to the previous version. A dedicated numerical forecasting branch is introduced, conditioned on multimodal representations from the LLM and time series encoder, preserving numerical fidelity while maintaining computational efficiency.
Pre-training:
-
Visual Pre-training (VP): Learns from large-scale unlabeled scientific documents rendered as page images, preserving figures, tables, equations, and layout information. It uses a contrastive next-latent prediction objective and requires neither OCR nor paired data.
-
Interleaved Text-Image Data: Constructed from PDFs using MinerU2.5-Pro for OCR and layout-aware structural parsing. Visual units (images, interline equations, tables) are cropped and reorganized into page-level sequences. A visual-gain-based quality filtering mechanism retains only pages where visual content significantly reduces perplexity.
-
Image Retrieval Enhancement: A large-scale image retrieval pipeline builds a vector database using an 8B embedding model. It supports text-to-image and image-to-image retrieval with post-processing including deduplication and reranking.
Post-Training:
-
Supervised Fine-Tuning: Converts the pretrained model into a controllable assistant using a large-scale, high-quality multimodal dataset covering general conversation, instruction following, code, image-text understanding, tool use, and scientific tasks.
-
Scalable and Stable Reinforcement Learning:
-
Partial Rollout with Off-Policy Correction: A co-located partial-rollout system pauses in-flight rollouts during policy updates, retaining prefixes and metadata. Importance-sampling ratios correct for policy mismatch, with clipping and token masks for numerical consistency.
-
Adaptive Length Regularization: Reweights advantages of positive responses to encourage concise reasoning only when the model has largely mastered a query, without penalizing negative responses.
-
Speculative Decoding: A draft model is trained online using the latest policy's token distributions, with a hybrid LK Loss combining forward KL divergence and total variation distance. This delivers 2× speedup in rollout generation and 1.7× end-to-end speedup.
-
Robust Multi-Task Optimization: Group-level Entropy-Controlled Policy Optimization (GEPO) uses group-level entropy to attenuate positive advantages in low-entropy groups and negative advantages in high-entropy groups, balancing exploration across heterogeneous tasks.
-
Unified RL Objective: Combines leave-one-out REINFORCE with GEPO, adaptive length regularization, clipped importance weights, and BKL masks.
-
Large-Scale Black- and White-Box Agentic RL:
-
A unified framework based on a harness × task abstraction decouples agent execution interfaces from task distributions.
-
Supports white-box and black-box harnesses (e.g., OpenClaw, Claude Code, OpenCode, OpenHands, Mini-SWE) through adapters.
-
A token-in–token-out (TITO) serving interface captures exact token IDs, log probabilities, and router experts.
-
Trace-aware experience assembly uses an incremental PrefixTree to align semantic trajectories with token-level evidence.
-
Tasks are constructed from public coding and terminal benchmarks plus a self-evolving task-synthesis system based on community skills, with stage-wise validation and step-level curation.
-
Process-aware advantage control applies penalties to specific assistant messages exhibiting deterministic process errors, while verifier integrity measures prevent reward hacking.
-
On-Policy Distillation: Consolidates separately optimized reasoning and agentic expert policies into a unified model. Uses a lightweight SFT warmup to reduce initial policy discrepancy, then maximizes negative reverse KL divergence with sampled-token teacher log-probabilities, reducing communication payload from O(HV) or O(Hk) to O(H).
Evaluation Results:
-
Scientific Benchmarks: Intern-S2-Preview-397B outperforms strong open- and closed-source models on Biology-Instructions (56.92), Mol-Instructions (52.37), and SciReasoner (63.97). It achieves state-of-the-art results on internal MP20 (67.88) and ProteinBinder-9 (4.36) evaluation sets. It delivers the best performance among open-source models on MolecularIQ (61.49), TOMG-Bench (65.66), XLRS-Bench (51.97), and MicroVQA (68.81).
-
General Benchmarks: Achieves the best results among open-source models on MMLU-Pro (89.75), SimpleQA-Verified (69.90), MMMU-Pro (80.46), and ChartQAPro (69.65). On general-purpose agentic tasks, it consistently outperforms Qwen3.5-397B and demonstrates performance comparable to Kimi-K2.7-Code.
-
Time Series Understanding: On SciTS, Intern-S2-Preview-397B achieves comparable or better performance than the trillion-parameter Intern-S1-Pro on seven of nine tasks, with improvements particularly pronounced on ASU03, BIU01, BIU03, MEU01, and PHU01. It also extends to radar coding-scheme classification and mode-and-modulation classification.
-
Time Series Generation: On SciTS forecasting tasks, Intern-S2-Preview-397B outperforms specialised time series baselines, with clear gains on ENG02, ENG03, MEG03, PHG02, and URG05. The horizon predictor achieves 99% accuracy. On GIFT-Eval, it achieves a competitive zero-shot MASE of 0.785.
-
Memory Decoder: The memory-augmented variant (Intern-MemDec-4B) improves the Biology-Instructions average score from 56.92 to 60.32 relative to the frozen backbone, while remaining close on cross-domain benchmarks, indicating targeted scientific specialization without compromising general capabilities.
Improvements for AI systems
Improvements to AI Systems:
- Dynamic Memory-Augmented Specialization without Retraining
-
Integrate a lightweight, token-level routed memory decoder alongside any frozen large backbone.
-
Use retrieval distillation to compress domain-specific evidence (e.g., biology, chemistry) into parametric memory.
-
At inference, fuse backbone and memory-decoder next-token distributions dynamically.
-
Resulting capability: Rapid, targeted specialization (e.g., +3.4 points on Biology-Instructions) without catastrophic forgetting or modifying the base model—ideal for multi-tenant deployments where one backbone serves many specialized domains.
- Stable Multi-Task Reinforcement Learning with Entropy Control
-
Apply group-level entropy monitoring to reweight advantages: suppress positive advantages in low-entropy (overconfident) groups and negative advantages in high-entropy (uncertain) groups.
-
Combine with leave-one-out REINFORCE and adaptive length regularization that only penalizes verbosity after task mastery.
-
Resulting capability: A single policy that balances exploration across heterogeneous tasks (e.g., coding, math, scientific QA) without mode collapse or reward hacking, improving worst-case task performance while maintaining peak performance.
- Partial Rollout with Off-Policy Correction for Efficient RL
-
Pause in-flight rollouts during policy updates; reuse prefixes and metadata with importance-sampling ratios (clipped, token-masked) to correct policy mismatch.
-
Resulting capability: Up to 2× faster RL training cycles with lower GPU memory, enabling larger batch sizes and more frequent policy updates—critical for scaling to 397B+ parameters.
- Online Speculative Decoding with Hybrid KL Loss
-
Train a draft model online using the latest policy’s token distributions, optimizing a hybrid loss (forward KL + total variation distance).
-
Resulting capability: 1.7× end-to-end inference speedup during RL rollout generation, reducing wall-clock time for agentic tasks and enabling real-time interactive scientific reasoning.
- Trace-Aware Experience Assembly for Agentic RL
-
Use an incremental PrefixTree to align semantic trajectories (e.g., tool calls, code execution) with token-level evidence (log-probs, router experts).
-
Apply process-aware advantage penalties to assistant messages with deterministic errors (e.g., syntax mistakes) and verifier integrity checks to prevent reward hacking.
-
Resulting capability: More reliable training signal for long-horizon agentic tasks (e.g., SWE-bench, terminal use), reducing spurious successes and improving task completion rates.
- On-Policy Distillation from Multiple Expert Policies
-
After separately optimizing reasoning and agentic policies, consolidate them into one model via lightweight SFT warmup + negative reverse KL divergence with sampled-token teacher log-probs.
-
Resulting capability: A unified model that retains both deep reasoning (scientific benchmarks) and tool-use/agentic skills without needing to switch between checkpoints—reducing serving costs and latency.
- Efficient Long-Sequence Time Series Encoder with Forecasting Branch
-
Use a multi-channel encoder supporting up to 300k time steps with 5–6× faster inference and 20% GPU memory.
-
Add a numerical forecasting branch conditioned on multimodal LLM representations.
-
Resulting capability: Real-time analysis and forecasting of extremely long sensor/physiological/radar signals (e.g., EEG, climate data) with 99% horizon-prediction accuracy, surpassing specialized time-series models.
- Visual-Gain-Based Data Filtering for Multimodal Pre-training
-
Retain only pages where visual content (figures, tables, equations) significantly reduces perplexity over text-only baselines.
-
Resulting capability: Higher-quality multimodal training data, reducing noise and improving downstream visual reasoning (e.g., +5–10% on chart and diagram QA) with fewer training steps.
- Unified Harness × Task Abstraction for Agentic RL
-
Decouple agent execution interfaces (white-box, black-box) from task distributions via adapters (e.g., OpenClaw, Claude Code, OpenHands).
-
Resulting capability: Plug-and-play training across diverse agent environments (coding, terminal, web) with minimal code changes, enabling rapid scaling to new benchmarks.
- Self-Evolving Task Synthesis with Stage-Wise Validation
-
Generate new agentic tasks from community skills, then validate and curate them in stages (syntax → semantic → execution).
-
Resulting capability: Continuous expansion of training task diversity without manual curation, improving generalization to unseen real-world tool-use scenarios.
What the Improved AI System Can Do:
-
Scientific Research Assistant: Specialize instantly in biology, chemistry, or materials science by attaching a memory decoder—no retraining needed. It can read papers, extract molecular properties, answer complex multi-hop questions (e.g., 60+ on Biology-Instructions), and forecast experimental outcomes from time-series sensor data.
-
Autonomous Coding & Terminal Agent: Execute long-horizon tasks (e.g., repository-level bug fixes, system administration) with high reliability, using trace-aware RL and process-error penalties. It matches or exceeds specialized code agents (e.g., Kimi-K2.7) while maintaining strong general reasoning (89.75 on MMLU-Pro).
-
Real-Time Scientific Monitoring System: Process 300k-step time series (e.g., EEG, radar, climate) 5–6× faster than previous models, with numerical forecasting accuracy (99% horizon prediction) and zero-shot generalization to new signal types (MASE 0.785 on GIFT-Eval).
-
Multi-Domain Unified Assistant: One frozen backbone serves many specialized domains (via memory decoders) and multiple agentic harnesses (via adapters), reducing serving cost and complexity while improving task-specific performance.
-
Efficient RL Training Infrastructure: Train 397B+ models with 2× faster rollouts, stable multi-task optimization, and on-policy distillation—enabling rapid iteration on new benchmarks without sacrificing performance or stability.
Abstract
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.
Sources
- GPT-4 Technical Report
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
- ClawGym: A Scalable Framework for Building Effective Claw Agents
- Intern-S1: A Scientific Multimodal Foundation Model
- SRT: Accelerating Reinforcement Learning via Speculative Rollout with Tree-Structured Cache
- Accelerating Large Language Model Decoding with Speculative Sampling
- ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving
- A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
- Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks