Very simplistic lipsyncing to get by. It uses three blend shapes for the lips: kiss, lips closed or mouth open. In terms of Facial Action Units: AU22, AU24 and AU27. It works with speech audio files or audio streams (no need to transcribe phonemes or text).
The technology is based on the differences between the spectrum of vowels and consonants. An explanation and demos of it can be seen in this video (link starting at the demos):
• Web-based live speech-driven lip-sync. An ...
The Unity package can be downloaded from here:
https://doi.org/10.5281/zenodo.5765691
Github (javascript version):
https://github.com/gerardllorach/thre...
More info and live demos:
https://gerardllorach.weebly.com/
References:
Llorach, G., Evans, A., Blat, J., Grimm, G., & Hohmann, V. (2016, September). Web-based live speech-driven lip-sync. In 2016 8th International Conference on Games and Virtual Worlds for Serious Applications (VS-GAMES) (pp. 1-4). IEEE.