Finance / Banking
China Merchants Bank: Shared operation of accelerators for AI training and inference
- Company
- China Merchants Bank
- Country
- China
- Adoption stage
- In operation
- Source published
- Date basis
- The date the source was published. It can differ from the date adoption started.
- How the source was checked
- Read the full source text
The work problem
As China Merchants Bank extended AI across more financial use cases, it had to run large model training, fine tuning and online inference together on nearly 10,000 heterogeneous accelerator cards. Adding more cards would not solve the problem. A training job can only start useful work once all the required workers and cards are ready, so it needs stable capacity, while online inference has to scale up and down quickly with request volume. Whenever several tenants fine tuned the same base model, a separate copy of that base model was loaded for each of them.
Technology and data
The bank built a single control plane on Kubernetes while keeping the execution paths for training and inference separate. Kueue first decides whether a training job can enter the cluster based on queues and quotas. The Kubernetes scheduler and HAMi then assign accelerator capacity, and Fluid speeds up access to datasets and model weights. Twinkle, developed in house, runs training on Ray, and vLLM or SGLang serves online inference. On the inference side Prometheus collects request rate, queue depth and latency, and KEDA uses those signals to add or remove replicas. Platform architects decided that training and inference needed separate admission policies.
Results
The bank said it brought 99 percent of its accelerator compute resources under this framework. Average accelerator compute utilization rose from 35 percent to more than 60 percent, and under comparable model and service conditions the inference cost of processing one million tokens fell by more than 60 percent. With Twinkle, five tenants fine tuning with LoRA share one base model instance. Base model replicas drop from five to one, accelerator resource usage falls by 80 percent and the number of tenants that can train at the same time increases fivefold. The bank wrote that it uses five tenants by default in production and has validated eight.
Limits and open questions
The source does not say when the platform went into operation or over what period each figure was measured. The numbers come from the bank platform team comparing the same internal measurement method before and after the redesign.
Sources
Compiled from public sources. These are not results from ATF Works customers.