Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Published
- Source
- arXiv
- Paper number
- 146
- Field
- LLMs / Architecture
- arXiv ID
- 2604.12374
Key points
- Existing large language models struggle with reasoning efficiency and real-world deployment cost, and earlier Mixture-of-Experts (MoE) designs often ignored real latency and memory-bandwidth constraints.
- Scaling large language model pretraining to trillions of tokens while preserving stability and accuracy in low-precision formats remains unsolved, and it hinders energy efficiency and hardware friendliness.
- Developing robust and generalizable agentic reasoning capabilities in LLMs for multi-stage tool use and long-horizon tasks requires both specialized architectural and training approaches.
- The paper introduces a hybrid Mamba-attention Mixture-of-Experts (MoE) architecture with a new LatentMoE design for efficient routing and a Multi-Token Prediction (MTP) mechanism that improves modeling quality and speculative decoding.
- It performs large-scale pretraining on 25 trillion tokens using NVIDIA's proprietary low-precision NVFP4 format, demonstrating stability and accuracy at scale.
- It implements a multi-stage post-training pipeline that combines extensive supervised fine-tuning (SFT) with multi-environment reinforcement learning from verifiable rewards (RLVR), with a focus on agentic reasoning and long-context capability up to 1M tokens.
Paper links
External research summaries. These are not HDATF publications or measured product results.