Fetching the paper…
Reading the bibliography…
Neural language modeling (LM) has led to significant improvements in several applications, including Automatic Speech Recognition.
Statistical phrase-based translation
Philipp Koehn, Franz Josef Och, and Daniel Marcu. 2003 · 2003
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou. 2010 · 2010
Earlier work this paper cites.
Variational approximation of long-span language models for lvcsr
Anoop Deoras, Tomáš Mikolov, Stefan Kombrink, Martin Karafiát, and Sanjeev Khudanpur. 2011 · 2011
Earlier work this paper cites.
Speech recognition and keyword spotting for low-resource languages: Babel project research at cued
Mark JF Gales, Kate M Knill, Anton Ragni, and Shakti P Rath. 2014 · 2014
Earlier work this paper cites.
Multi-task learning for multiple language translation
Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, and Haifeng Wang. 2015 · 2015
Earlier work this paper cites.
Massively multilingual word embeddings
Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, and Noah A Smith. 2016 · 2016
Earlier work this paper cites.
Jointly learning to embed and predict with multiple languages
Daniel C. Ferreira, André F. T. Martins, and Mariana S. C. Almeida. 2016 · 2016
Earlier work this paper cites.
Multi-way, multilingual neural machine translation with a shared attention mechanism
Orhan Firat, Kyunghyun Cho, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Multi-language neural network language models
A Ragni, E Dakin, X Chen, MJF Gales, and KM Knill. 2016 · 2016
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson et al. 2017 · 2017
Cited alongside, same era.
Multilingual hierarchical attention networks for document classification
Nikolaos Pappas and Andrei Popescu-Belis. 2017 · 2017
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Later among the works it cites.
Retrieve-and-read: Multi-task learning of information retrieval and reading comprehension
Kyosuke Nishida, Itsumi Saito, Atsushi Otsuka, Hisako Asano, and Junji Tomita. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Self-training for jointly learning to ask and answer questions
Mrinmaya Sachan and Eric Xing. 2018 · 2018
Later among the works it cites.
Neural document summarization by jointly learning to score and select sentences
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. 2018 · 2018
Later among the works it cites.