Fetching the paper…
Reading the bibliography…
Sequence-to-sequence (Seq2Seq) models with attention have excelled at tasks which involve generating natural language sentences such as machine translation, image captioning and speech recognition.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Statistical Language Models Based on Neural Networks
Mikolov, T · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, KyungHyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Amodei, Dario, Anubhai, Rishita, Battenberg, Eric, Case, Carl, Casper, Jared, Catanzaro, Bryan, Chen, Jingdong, Chrzanowski, Mike, Coates, Adam, Diamos, Greg, et al · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, Samy, Vinyals, Oriol, Jaitly, Navdeep, and Shazeer, Noam · 2015
Earlier work this paper cites.
Chan, William, Jaitly, Navdeep, Le, Quoc V, and Vinyals, Oriol · 2015
Cited alongside, same era.
On using monolingual corpora in neural machine translation
Gulcehre, Caglar, Firat, Orhan, Xu, Kelvin, Cho, Kyunghyun, Barrault, Loic, Lin, Huei-Chi, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2015
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Sennrich, Rico, Haddow, Barry, and Birch, Alexandra · 2015
Cited alongside, same era.
Towards better decoding and language model integration in sequence to sequence models
Chorowski, Jan and Jaitly, Navdeep · 2016
Later among the works it cites.
Exploring the limits of language modeling
Jozefowicz, Rafal, Vinyals, Oriol, Schuster, Mike, Shazeer, Noam, and Wu, Yonghui · 2016
Later among the works it cites.
Unsupervised pretraining for sequence to sequence learning
Ramachandran, Prajit, Liu, Peter J, and Le, Quoc V · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Yonghui, Schuster, Mike, Chen, Zhifeng, Le, Quoc V, Norouzi, Mohammad, Macherey, Wolfgang, Krikun, Maxim, Cao, Yuan, Gao, Qin, Macherey, Klaus, et al · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vinyals, Oriol and Le, Quoc V · 2015
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
Bahdanau, Dzmitry, Chorowski, Jan, Serdyuk, Dmitriy, Brakel, Philemon, and Bengio, Yoshua · 2016
Cited alongside, same era.
Attention-based models for speech recognition
Chorowski, Jan, Bahdanau, Dzmitry, Serdyuk, Dmitry, Cho, Kyunghyun, and Bengio, Yoshua
Cited in the paper.
Attention-based models for speech recognition
Chorowski, Jan K, Bahdanau, Dzmitry, Serdyuk, Dmitriy, Cho, Kyunghyun, and Bengio, Yoshua
Cited in the paper.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V
Cited in the paper.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V
Cited in the paper.
Yang, Zhilin, Dhingra, Bhuwan, Yuan, Ye, Hu, Junjie, Cohen, William W, and Salakhutdinov, Ruslan · 2016
Later among the works it cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, Noam, Mirhoseini, Azalia, Maziarz, Krzysztof, Davis, Andy, Le, Quoc, Hinton, Geoffrey, and Dean, Jeff · 2017
Closest in time.