This talk will focus on high-performance and scalable middleware for distributed deep learning training, machine learning and inference for various AI workload on modern heterogeneous clusters. On the distributed training side, we will focus on: MPI-driven solutions (MPI4DL and Mix-and-Match Runtime (MCL-DL)) to extract performance and scalability for PyTorch and TensorFlow-based frameworks with different parallelism (data, model, layer, pipeline, and spatial). On the machine learning side, we will focus on MPI-driven ML training with MPI4cuML. On the inference side, we will present novel parallel inference techniques (ParaInfer-X) to facilitate deployment of emerging AI models on edge devices and HPC clusters.