Speaker: Daksh Varshneya, Rasa
Host: Nishant Sinha, OffNote Labs
Leveraging real dialogues for better task-oriented assistants
Recently, there has been a steady increase in the number of dialogue datasets available in the open source domain. This presents a good opportunity for the dialogue community to build better pre-trained models which can efficiently encode conversational data.
In this talk, we first try to understand why pre-training on conversational data should help dialogue systems. Next, we deep dive into a state of the art NLU architecture DIET which is designed for efficient intent classification and entity recognition. Further, we'll see how data augmentation through paraphrasing can help improve DIET's performance. Lastly, we'll introduce the concept of Conversation Driven Development(CDD) and see how it helps to improve your assistant over time by learning from data coming from real users.
===
The OffNote Labs AI Talk Series brings you industry experts, researchers and practitioners, passionate to share their learnings and experience -- on innovating and building cutting-edge AI technology / systems which touch and influence the lives of billions of people on this planet.
Our talks are informal, a blend of traditional presentation / podcast, and the audience very technically engaged.
We take delight in unraveling the experience of research and innovation, and celebrate innovators who 'follow the problem', and make complex technology work in the wild.
==
Follow OffNote Labs on LinkedIn. / offnote
Watch previous talks and Subscribe on Youtube: https://bit.ly/31VTMHH
Sign up to our newsletter: http://offnote.substack.com
Read our research articles: / offnote
Web: https://offnote.co
===
00:00 Introduction - Offnote Labs
14:51 Overview - Towards Better Task-Oriented Assistants
16:15 Introduction: Speaker Dash Varshneya, Rasa
17:55 AI Assistants
19:01 Modular Task-Oriented Assistants - NLU
25:25 Modular Task-Oriented Assistants - Dialogue Management
28:11 Evaluate - Pre Trained Encoder
28:45 Representations to Capture
32:21 Conversational data v/s Prose data
35:58 A Dataset for Evaluation
36:48 Encoders to Evaluate
44:01 Conversational Data Encodes Conversational Cues
44:35 Use Unsupervised Learning for Intrinsic Investigation
48:28 Story for Extrinsic Evaluation
54:51 A Dataset for Evaluation
55:08 Unsupervised Learning for Intrinsic Investigation
59:35 Pre Trained Encoders: Effective Choice for Featurization
1:00:08 Revisit: NLU Tasks
1:02:38 Lightweight Language Understanding for Dialogue Systems
1:08:28 Sparse + Conversational Embeddings
1:22:05 Data Augmentation
1:22:21 Build an Assistant - Not Easy
1:23:28 Not to Paraphrase
1:24:51 Multilingual Neural Machine Translation
1:26:15 Multilingual NMT Models: Paraphrasing
1:28:45 Beam Search
1:30:41 Multilingual NMT+ DBS work
1:31:31 Improved Performance on Intent Classification
1:43:45 Review, Annotate, Test, Track and Fix
1:47:21 Not a Linear Process: Jumping between actions
1:48:21 Conversation Driven Development(CDD), Putting CDD to test
1:49:51 Questions and Conclusion
#Offnote #offnotetalks #offnotelabs #taskoriented #modulartask #dialougemanagement #pretrained #encoder #NLU #Prosedata #conversationaldata #extrinsicevaluation #multilingual #DBS #NMT #dataaugmentation #beamsearch #annotate #translation #unsupervisedlearning #dialougedataset #CDD #DIET #conversationdrivendevelopment