In this video, I show you how to accelerate Transformer inference with AWS Inferentia, a custom chip designed by AWS.
Starting from a BERT model that I fine-tuned on AWS Trainium ( • Accelerate Transformer training with ... ) , I compile it with the Neuron SDK for Inferentia. Then, using an inf1.6xlarge instance (4 Inferentia chips, 16 Neuron Cores), I show you how to use pipeline mode to predict at scale, reaching over 4,000 predictions per second at 3-millisecond latency.
⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos ⭐️⭐️⭐️
⭐️⭐️⭐️ Want to buy me a coffee? I can always use more :) https://www.buymeacoffee.com/julsimon ⭐️⭐️⭐️
Amazon EC2 Inf1: https://aws.amazon.com/ec2/instance-t...
AWS Neuron SDK documentation: https://awsdocs-neuron.readthedocs-ho...
AWS blog post: https://aws.amazon.com/fr/blogs/machi...
Setup steps and code: https://gitlab.com/juliensimon/huggin...
Interested in hardware acceleration for Transformers? Check out my other videos :
Training on Habana Gaudi: • Accelerate Transformer training with ...
Training on Graphcore: • Accelerate Transformer training with ...
Predicting with ONNX: • Accelerate Transformer inference on C...
Predicting with Intel OpenVINO: • Accelerate Transformer inference on C...
Inferentia compilation on SageMaker: • Accelerate Transformers on Amazon Sag...