Accelerating CPU-Based Deep Learning Training

X-Scale-AI-DPU

Our high performance AI solution leverages HPC technology to accelerate CPU-based distributed DNN training by utilizing the capabilities of DPUs.

Efficient Distributed Training

Distributed training with Pytorch using Horovod, along with enabling fine-tuned MPI library for CPU and DPU systems.

Rigorous Testing

Our team of expert software engineers tested this on several DNNs and datasets with up to 19% improvement in DNN training performance without checkpointing. With checkpointing, up to 33% improvement in epoch time.

Simplified Approach

Our user-friendly approach starts at installation with install and execution in one command.
Additionally, the user-friendly Python interface enables DL applications to run on the CPU and DPU.

Integrated Expert Support

Support for DNN checkpointing with DPUs and BlueField-3 DPUs, with support for more system configurations continually being developed.

Optimal Performance

Optimal Performance Offering “out-of-the-box” optimal performance on CPU+DPU platforms.

Performance Case Studies

Accelerating DNN Training Across Configurationser

Questions on X-ScaleAI-DPU?

Our team of experts are available to answer your questions and show you an X-ScaleAI-DPU demo to help you accelerate deep learning training for your CPU-based systems.