Accelerate Transformer inference with AWS Inferentia

Опубликовано: 25 Октябрь 2024
на канале: Julien Simon
2,483
39

In this video, I show you how to accelerate Transformer inference with AWS Inferentia, a custom chip designed by AWS.

Starting from a BERT model that I fine-tuned on AWS Trainium (   • Accelerate Transformer training with ...  ) , I compile it with the Neuron SDK for Inferentia. Then, using an inf1.6xlarge instance (4 Inferentia chips, 16 Neuron Cores), I show you how to use pipeline mode to predict at scale, reaching over 4,000 predictions per second at 3-millisecond latency.

⭐️⭐️⭐️ Don't forget to subscribe to be notified of future videos ⭐️⭐️⭐️
⭐️⭐️⭐️ Want to buy me a coffee? I can always use more :) https://www.buymeacoffee.com/julsimon ⭐️⭐️⭐️

Amazon EC2 Inf1: https://aws.amazon.com/ec2/instance-t...
AWS Neuron SDK documentation: https://awsdocs-neuron.readthedocs-ho...
AWS blog post: https://aws.amazon.com/fr/blogs/machi...
Setup steps and code: https://gitlab.com/juliensimon/huggin...

Interested in hardware acceleration for Transformers? Check out my other videos :
Training on Habana Gaudi:    • Accelerate Transformer training with ...  
Training on Graphcore:    • Accelerate Transformer training with ...  
Predicting with ONNX:    • Accelerate Transformer inference on C...  
Predicting with Intel OpenVINO:    • Accelerate Transformer inference on C...  
Inferentia compilation on SageMaker:    • Accelerate Transformers on Amazon Sag...