Shanghai AI Laboratory has introduced Intern-S2-397B, a 403-billion-parameter vision-language model for long-horizon scientific research and scientific agents. The lab announced the model in a post on X.
The lab said the model has capabilities in knowledge, coding and agent work, and ranks among the top open-source models. That is the lab's own assessment.
Its model card says vision-language pre-training handles scientific document pages and visual reasoning directly, without an intermediate parsing step. The model was trained on scientific reinforcement-learning tasks across more than 20 domains, including biomolecular interaction design and material structure generation.
The model supports long-horizon agent tasks through black-box agentic reinforcement learning and by connecting multiple agent frameworks to large-scale sandboxed environments. Its thinking mode is enabled by default, and it supports tool calling.
Intern-S2-397B is available in Safetensors with BF16 and F32 tensor types under the Apache 2.0 license. The Hugging Face model page lists a 256K-token text context window and a 64K-token multimodal context window, along with vLLM, SGLang and LMDeploy deployment support. The GitHub repository also lists BF16 and FP8 model formats.
Shanghai AI Laboratory is a national-level, state-backed research institute officially unveiled at the World Artificial Intelligence Conference in July 2020. Its INTERN large-model system predates this release: the original InternLM was jointly developed with SenseTime, the Chinese University of Hong Kong and Fudan University, followed by InternLM2 in January 2024, InternLM2.5 in July 2024 and InternLM3 in January 2025.




