Deep Neural Networks (DNNs) are powerful models that have achieved excellent performance on difficult learning tasks. Although DNNs work well whenever large labeled training sets are available, they cannot be used to map sequences to sequences. In this paper, we present a general end-to-end approach to sequence learning that makes minimal assumptions on the sequence structure. Our method uses a multilayered Long Short-Term Memory (LSTM) to map the input sequence to a vector of a fixed dimensionality, and then another deep LSTM to decode the target sequence from the vector. Our main result is that on an English to French translation task from the WMT'14 dataset, the translations produced by the LSTM achieve a BLEU score of 34.8 on the entire test set, where the LSTM's BLEU score was penalized on out-of-vocabulary words. Additionally, the…
PAPER
Sequence to Sequence Learning with Neural Networks
http://arxiv.org/abs/1409.3215
LISTEN
https://researchpod.app/episode/d38be...
https://researchpod.app
AUTHORS
Ilya Sutskever, Oriol Vinyals, Quoc V. Le
TOPICS
Computer science, Artificial intelligence, Sentence, Phrase, Sequence (biology)
ABOUT RESEARCHPOD
ResearchPod turns research papers into podcast episodes so you can keep up with science while listening.
https://researchpod.app
DISCLAIMER
This is an AI-generated podcast discussion of the paper and is not a substitute for reading the original work.