Natural oral conversation is the Holy Grail of conversational AI, how close are we?
Historically large language models have been exclusively text based at their core. Pipelines for "talking" to them have involved separate Speech to Text, Inferencing, and Speech to Text steps. This causes latency, errors and loss of audio only nuance in the STT path. We can use various techniques to try and mitigate around these, but the results are, at best, sub-optimal.
This talk explores the state of the art in technical detail and analyses whether tokenisation of audio directly into LLMs will deliver more natural conversational results and where we are at with this.
---------
Learn more about RTC.ON:
▶️ https://rtcon.live/
Follow us on X:
▶️ https://x.com/ElixirMembrane
▶️ https://x.com/swmansion
#ai #streaming #webrtc