Fetching the paper…
Reading the bibliography…
Recently, substantial progress has been made in language modeling by using deep neural networks.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Support vector machines for multi-class pattern recognition
Jason Weston, Chris Watkins, et al · 1999
Earlier work this paper cites.
Large margin methods for structured and interdependent output variables
Ioannis Tsochantaridis, Thorsten Joachims, Thomas Hofmann, and Yasemin Altun · 2005
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Reducing overfitting in deep networks by decorrelating representations
Michael Cogswell, Faruk Ahmed, Ross Girshick, Larry Zitnick, and Dhruv Batra · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2016
Earlier work this paper cites.
Statistics of robust optimization: A generalized empirical likelihood approach
John Duchi, Peter Glynn, and Hongseok Namkoong · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Large-margin softmax loss for convolutional neural networks
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Cited alongside, same era.
Convolutional sequence modeling revisited, 2018
Shaojie Bai, J. Zico Kolter, and Vladlen Koltun · 2018
Later among the works it cites.
Large margin deep networks for classification
Gamaleldin F Elsayed, Dilip Krishnan, Hossein Mobahi, Kevin Regan, and Samy Bengio · 2018
Later among the works it cites.
Frage: frequency-agnostic word representation
Chengyue Gong, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2018
Later among the works it cites.
Predicting the generalization gap in deep networks with margin distributions
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio · 2018
Later among the works it cites.
Sigsoftmax: Reanalysis of the softmax bottleneck
Sekitoshi Kanai, Yasuhiro Fujiwara, Yuki Yamanaka, and Shuichi Adachi · 2018
Later among the works it cites.
A la carte embedding: Cheap but effective induction of semantic feature vectors
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Cited alongside, same era.
AUTOMATIC SPEECH RECOGNITION
Dong Yu and Li Deng · 2016
Cited alongside, same era.
An actor-critic algorithm for sequence prediction
Dzmitry Bahdanau, Philemon Brakel, Kelvin Xu, Anirudh Goyal, Ryan Lowe, Joelle Pineau, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Quasi-recurrent neural networks
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
Noisy softmax: Improving the generalization ability of dcnn via postponing the early softmax saturation
Binghui Chen, Weihong Deng, and Junping Du · 2017
Cited alongside, same era.
Efficient softmax approximation for gpus
Edouard Grave, Armand Joulin, Moustapha Cissé, David Grangier, and Hervé Jégou · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Cited alongside, same era.
Mikhail Khodak, Nikunj Saunshi, Yingyu Liang, Tengyu Ma, Brandon Stewart, and Sanjeev Arora · 2018
Later among the works it cites.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2018
Later among the works it cites.
Robust statistics: theory and methods (with R)
Ricardo A Maronna, R Douglas Martin, Victor J Yohai, and Matías Salibián-Barrera · 2018
Later among the works it cites.
All-but-the-top: Simple and effective postprocessing for word representations
Jiaqi Mu, Suma Bhat, and Pramod Viswanath · 2018
Later among the works it cites.
Fast parametric learning with activation memorization
Jack W Rae, Chris Dyer, Peter Dayan, and Timothy P Lillicrap · 2018
Later among the works it cites.
Regularizing deep networks using efficient layerwise adversarial training
Swami Sankaranarayanan, Arpit Jain, Rama Chellappa, and Ser Nam Lim · 2018
Later among the works it cites.
Tensor2tensor for neural machine translation
Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N. Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, Ryan Sepassi, Noam Shazeer, and Jakob Uszkoreit · 2018
Later among the works it cites.
Additive margin softmax for face verification
Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu · 2018
Later among the works it cites.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2018
Later among the works it cites.
Representation degeneration problem in training natural language generation models
Jun Gao, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tieyan Liu · 2019
Closest in time.
Sentence-wise smooth regularization for sequence to sequence learning
Chengyue Gong, Xu Tan, Di He, and Tao Qin · 2019
Closest in time.
Partially shuffling the training data to improve language models
Ofir Press · 2019
Closest in time.
Multi-agent dual learning
Yiren Wang, Yingce Xia, Tianyu He, Fei Tian, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu · 2019
Closest in time.