Inference & GPU Optimization: GPTQ

Опубликовано: 27 Октябрь 2024
на канале: AI Makerspace
317
24

Join us for Part II of our series on optimizing inference for GPT models, where we'll explore the powerful Generative Pretrained Transformer Quantization (GPTQ) method. Building on our discussion of Activation-aware Weight Quantization (AWQ) from Part I, we’ll break down how GPTQ speeds up inference time, reduces latency, and right-sizes your model for better performance. Learn how GPTQ stacks up against AWQ, and get hands-on with code to understand its unique approach using second-order information. Optimize your LLMs like never before with cutting-edge quantization techniques!

Event page: https://bit.ly/InferenceGPTQ?utm_sour...

Have a question for a speaker? Drop them here:
https://app.sli.do/event/mZ4GbFrJd1Cm...

Speakers:
​Dr. Greg, Co-Founder & CEO AI Makerspace
  / gregloughane  

The Wiz, Co-Founder & CTO AI Makerspace
  / csalexiuk  

Apply for our new AI Engineering Bootcamp on Maven today!
https://bit.ly/aie1

For team leaders, check out!
https://aimakerspace.io/gen-ai-upskil...

Join our community to start building, shipping, and sharing with us today!
  / discord  

How'd we do? Share your feedback and suggestions for future events.
https://forms.gle/krsiK5132XaL7qPK8