Quantization and Llama 3.2 on mobile. Post-Training quantization SpinQuant

Опубликовано: 04 Август 2026
на канале: Ruslan Dev
1,768
81

Subscribe to the Telegram channel: https://t.me/ruslandevlive

This video is about Llama 3.2 quantization and its optimization for mobile devices, as well as about the new SpinQuant post-training quantization method.

💻 immers.cloud – a wide selection of cards for training and inference of neural networks: https://immers.cloud/signup/r/2024042...
One of the leading IaaS (Infrastructure as a Service) service providers in Russia, specializing in the use of graphics processing units (GPUs).
The service offers competitive prices and an intuitive interface that even novice users can easily master and start working with the necessary software.

💻 gptchain – a framework for quickly deploying AI assistants: https://github.com/RuslanPeresy/gptchain
Supports integration with a Telegram bot, Retrieval Augmented Generation (RAG), deploying models to an LLM server, and fine-tuning LLM on your own data.

Discord:   / discord  

This description contains referral links.