Fetching the paper…
Reading the bibliography…
In this paper we show that reporting a single performance score is insufficient to compare non-deterministic approaches.
The kolmogorov-smirnov test for goodness of fit
Frank J. Massey. 1951 · 1951
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O(1/sqr(k))
Yurii Nesterov. 1983 · 1983
Earlier work this paper cites.
Backpropagation Applied to Handwritten Zip Code Recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel. 1989 · 1989
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Efficient BackProp
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller. 1998 · 1996
Earlier work this paper cites.
Feature-rich Part-of-speech Tagging with a Cyclic Dependency Network
Kristina Toutanova, Dan Klein, Christopher D. Manning, and Yoram Singer. 2003 · 2003
Earlier work this paper cites.
Why Does Unsupervised Pre-training Help Deep Learning?
Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio. 2010 · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Neural Networks for Machine Learning - Lecture 6a - Overview of mini-batch gradient descent
Geoffrey Hinton. 2012 · 2012
Earlier work this paper cites.
Statistical language models based on neural networks
Tomáš Mikolov. 2012 · 2012
Cited alongside, same era.
ADADELTA: an adaptive learning rate method
Matthew D. Zeiler. 2012 · 2012
Cited alongside, same era.
Offspring from Reproduction Problems: What Replication Failure Teaches Us
Antske Fokkens, Marieke van Erp, Marten Postma, Ted Pedersen, Piek Vossen, and Nuno Freire. 2013 · 2013
Cited alongside, same era.
Joint Event Extraction via Structured Prediction with Global Features
Qi Li, Heng Ji, and Liang Huang. 2013 · 2013
Cited alongside, same era.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Bidirectional LSTM-CRF Models for Sequence Tagging
Zhiheng Huang, Wei Xu, and Kai Yu. 2015 · 2015
Later among the works it cites.
Globally normalized transition-based neural networks
Daniel Andor, Chris Alberti, David Weiss, Aliaksei Severyn, Alessandro Presta, Kuzman Ganchev, Slav Petrov, and Michael Collins. 2016 · 2016
Later among the works it cites.
Enriching Word Vectors with Subword Information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2016 · 2016
Later among the works it cites.
A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Later among the works it cites.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. 2013 · 2013
Cited alongside, same era.
An Efficient Approach for Assessing Hyperparameter Importance
Frank Hutter, Holger Hoos, and Kevin Leyton-Brown. 2014 · 2014
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Dependency-Based Word Embeddings
Omer Levy and Yoav Goldberg. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
Incorporating Nesterov Momentum into Adam
Timothy Dozat. 2015 · 2015
Cited alongside, same era.
Flat Minima
Sepp Hochreiter and Jürgen Schmidhuber. 1997a
Cited in the paper.
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Later among the works it cites.
Dependency based embeddings for sentence classification tasks
Alexandros Komninos and Suresh Manandhar. 2016 · 2016
Later among the works it cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Later among the works it cites.
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard H. Hovy. 2016 · 2016
Later among the works it cites.
Deep multi-task learning with low level tasks supervised at lower layers
Anders Søgaard and Yoav Goldberg. 2016 · 2016
Later among the works it cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Later among the works it cites.
Optimal Hyperparameters for Deep LSTM-Networks for Sequence Labeling Tasks
Nils Reimers and Iryna Gurevych. 2017 · 2017
Closest in time.