SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Published
Source
arXiv
Paper number
694
Field
LLMs / NLP
arXiv ID
2607.20145

Key points

  • On Ascend SuperPOD, it raises the MFU for DeepSeek-V4-Pro, a 1.6T model, from 11.67 percent to 34.22 percent, or 2.93x.
  • It uses an agent called AuraKernel to automatically optimize complex NPU kernels.
  • It builds a 10K high-quality CPT plus SFT dataset specialized for operations research and applies solver verification.
  • The specialized model achieves an average zero-shot Pass@1 of 71.81 percent, which is 3.98 percentage points above GPT-5.4-Mini and 11.27 points above the base model.
  • It provides a full-stack practical example of training a trillion-parameter model on NPU hardware rather than GPUs.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)