Text-to-speech (TTS) technologies have become an integral part of our daily lives, powering applications like GPS systems and voice assistants. However, a significant limitation is their difficulty in handling multilingual sentences, resulting in inaccurate pronunciation. This thesis aims to explore potential ways to address the challenge of word-level language detection which would allow to build TTS systems capable of handling these situations.