Large language models (LLMs) are useful for many user applications, from language translation to content creation, virtual assistance, and even cybersecurity. However, they are challenging to run on device. Take as an example Llama 2-7B, an LLM with 7 billion parameters. All the parameters must be read to generate each token, which equates to significant bandwidth, especially for providing long answers. As a result, memory bandwidth is the bottleneck for LLMs. Our research addresses this issue. In this NeurIPS 2023 demo, we show Llama 2-7B Chat running at a high token rate completely on device, both a smartphone and laptop, through full-stack AI optimization. We use speculative decoding and quantization-aware training with knowledge distillation.
Visit the Qualcomm AI Research website: https://www.qualcomm.com/research/art...
Develop with the Qualcomm AI Stack
https://www.qualcomm.com/products/tec...
Sign up for our newsletter
https://assets.qualcomm.com/mobile-co...