Inference & GPU Optimization: AWQ

Опубликовано: 28 Июль 2026
на канале: AI Makerspace
620
34

Join us as we explore cutting-edge techniques to optimize Large Language Models (LLMs) for inference! This event will dive into the tradeoffs between performance and cost in both LLMs and Small Language Models (SLMs). Learn how quantization, specifically Activate-aware Quantization (AWQ), compresses models while maintaining top-notch performance. We'll break down the findings from recent research and show you how to apply these techniques using Transformers. If you're interested in maximizing output while minimizing compute, this is an event you won't want to miss!

Join us every Wednesday at 1pm EST for our live events. SUBSCRIBE NOW to get notified!

Speakers:
​Dr. Greg, Co-Founder & CEO AI Makerspace
  / gregloughane  

The Wiz, Co-Founder & CTO AI Makerspace
  / csalexiuk  

Apply for The AI Engineering Bootcamp today!
https://bit.ly/AIEbootcamp

LLM Foundations - Email-based course
https://bit.ly/4izQVKk

Full LLM Engineering Cohort - Open-Source!
https://bit.ly/42ha7Xy

For team leaders, check out!
https://bit.ly/41YtlA7

Join our community to start building, shipping, and sharing with us today!
  / discord  

How'd we do? Share your feedback and suggestions for future events.
https://forms.gle/z96cKbg3epXXqwtG6

00:00:00 Introduction to GPU Optimization
00:03:39 Understanding GPU Constraints in LLMs
00:07:34 Understanding Low-Rank Adaptation (LoRA)
00:11:34 Understanding Quantization in Neural Networks
00:15:45 Focus on Posttraining Quantization (PTQ)
00:19:10 Introduction to Efficient Model Quantization
00:22:59 Enhancing Computing Efficiency with Kernel Fusion and Dequantization
00:27:00 GPU Optimization Techniques
00:30:42 Optimizing Performance with AWQ Tools
00:34:07 Understanding User Experience Through Token Latency
00:38:00 Optimizing Batch Size for Performance
00:41:42 Understanding Quantization Configurations and Auto AWQ
00:45:37 Accelerating Machine Learning Models with Quantization
00:49:12 Understanding Quantization and Scaling in Neural Networks
00:52:40 Exploring Smart Compression and Kernel Fusion
00:56:54 Drawbacks and Considerations in Using AWQ
00:59:45 Closing Remarks and Weekly Encouragement