SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
- Published
- Source
- arXiv
- Paper number
- 694
- Field
- LLMs / NLP
- arXiv ID
- 2607.20145
Key points
- On Ascend SuperPOD, it raises the MFU for DeepSeek-V4-Pro, a 1.6T model, from 11.67 percent to 34.22 percent, or 2.93x.
- It uses an agent called AuraKernel to automatically optimize complex NPU kernels.
- It builds a 10K high-quality CPT plus SFT dataset specialized for operations research and applies solver verification.
- The specialized model achieves an average zero-shot Pass@1 of 71.81 percent, which is 3.98 percentage points above GPT-5.4-Mini and 11.27 points above the base model.
- It provides a full-stack practical example of training a trillion-parameter model on NPU hardware rather than GPUs.
Paper links
External research summaries. These are not HDATF publications or measured product results.