Information technology / Translation software
DeepL: Improving the efficiency of training and running the translation models
- Company
- DeepL
- Country
- Germany
- Adoption stage
- In operation
- Source published
- Date basis
- The date the source was published. It can differ from the date adoption started.
- How the source was checked
- Read the full source text
The work problem
DeepL said that bringing in a DGX SuperPOD built from 544 NVIDIA H100 GPUs gave it a large increase in compute, but that while calculations ran in the 16 bit BF16 format the size of model it could grow within a given latency window and the number of requests it could take at once stayed tied together. The company explained that most of the compute in modern language models takes the form of matrix multiplication, so dropping the precision to 8 bits can raise throughput, but that FP8 carries half the range and precision of BF16 and therefore risks lowering training quality.
Technology and data
The company moved its existing training code from BF16 to FP8 using the NVIDIA Transformer Engine. Following NVIDIA's recommendation it uses the default setup, applying the more precise E4M3 in the forward pass that predicts the probability distribution of the next token, and the wider range E5M2 in the backward pass that computes gradients. It stores 32 bit scaling factors alongside the FP8 weight tensors to prevent overflow and underflow, and accounts for those factors when tensors are multiplied. After pre-training it fine tunes on specific tasks, distils large models into smaller ones, runs reinforcement learning and applies several parallelization strategies to use the large number of GPUs. For inference, NVIDIA TensorRT LLM builds an engine from the trained weights and applies kernel fusion, optimized CUDA code, KV caching and continuous in flight batching. Quality was checked by training a 1.5 billion parameter model on three trillion tokens in both formats and comparing training loss and validation perplexity for English and German.
Results
DeepL said model FLOPS utilization, which measures how much of the available compute the training actually uses, rose from 44.6% to 67% with FP8, making training 50% faster. On another training setup it said it worked with NVIDIA to refine its use of Transformer Engine features and added a further 25% over fifteen months, reaching 80% utilization. On quality it said training loss was very slightly lower for BF16, but the gap was drowned out by the step to step fluctuation in both formats, and it found no degradation in validation perplexity for English and German. In inference it said FP8 handled double the throughput of BF16 at the same latency for most batch sizes. The company said the result is that it can build models with far more parameters that translate 1.4 times better than its previous models for European languages and 1.7 times better for harder pairs such as English and Japanese, while still fitting inside the same latency window for production inference. It said language experts judged the quality.
Limits and open questions
The source does not say what evaluation method or how many comparisons produced the figures of 1.4 times and 1.7 times against the previous models. It also does not say when FP8 was first applied to production inference, nor how many language experts judged translation quality or against what criteria.
Sources
Compiled from public sources. These are not results from ATF Works customers.