Microsoft released an artificial intelligence tool named as VALL-E that can replicate people's voices just by listening 3-seconds audio of their speech.
Research Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers https://arxiv.org/abs/2301.02111
Source: https://valle-demo.github.io/