Qualcomm AI Research has used generative AI to develop a digital fitness coach that improves upon existing solutions in terms of accuracy and realism. The fitness coach provides real-time interaction by encouraging, correcting, and helping the user meet their fitness goals. Our demo showcases how a visually grounded large language model, or LLM, can enable natural interactions that are contextual, multimodal, and real-time. A video stream of the user exercising is processed by our action recognition model. Based on the recognized action, our stateful orchestrator grounds the prompt and feeds it to the LLM. The fitness coach provides the LLM answer back to the user through a text-to-speech avatar. This is made possible thanks to three key innovations: a vision model that is trained to detect fine-grained fitness activities, a language model that is trained to generate language grounded in the visual concepts, and an orchestrator that coordinates the fluid interaction between these two modalities to facilitate live dialogue coaching feedback. The result is a fitness coach that provides real-time interaction for an engaging and dynamic user experience.
Learn about our other CVPR 2023 activities: https://www.qualcomm.com/news/onq/202...
Visit the Qualcomm AI Research website: https://www.qualcomm.com/research/art...
Develop with the Qualcomm AI Stack
https://www.qualcomm.com/products/tec...