Join us for Part II of our series on optimizing inference for GPT models, where we'll explore the powerful Generative Pretrained Transformer Quantization (GPTQ) method. Building on our discussion of Activation-aware Weight Quantization (AWQ) from Part I, we’ll break down how GPTQ speeds up inference time, reduces latency, and right-sizes your model for better performance. Learn how GPTQ stacks up against AWQ, and get hands-on with code to understand its unique approach using second-order information. Optimize your LLMs like never before with cutting-edge quantization techniques!
Event page: https://bit.ly/InferenceGPTQ?utm_sour...
Have a question for a speaker? Drop them here:
https://app.sli.do/event/mZ4GbFrJd1Cm...
Speakers:
Dr. Greg, Co-Founder & CEO AI Makerspace
/ gregloughane
The Wiz, Co-Founder & CTO AI Makerspace
/ csalexiuk
Apply for our new AI Engineering Bootcamp on Maven today!
https://bit.ly/aie1
For team leaders, check out!
https://aimakerspace.io/gen-ai-upskil...
Join our community to start building, shipping, and sharing with us today!
/ discord
How'd we do? Share your feedback and suggestions for future events.
https://forms.gle/krsiK5132XaL7qPK8