Fetching the paper…
Reading the bibliography…
Top-performing machine learning systems, such as deep neural networks, large ensembles and complex probabilistic graphical models, can be expensive to store, slow to evaluate and hard to integrate into larger systems.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
Parallel distributed processing: Explorations in the microstructure of cognition
P. Smolensky · 1986
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
S. Becker and Y. Le Cun · 1989
Earlier work this paper cites.
Optimal brain damage
Y. Le Cun, J. S. Denker, and S. A. Solla · 1990
Earlier work this paper cites.
The evidence framework applied to classification networks
D. J. C. MacKay · 1992
Earlier work this paper cites.
Tangent prop—a formalism for specifying selected invariances in an adaptive network
P. Simard, B. Victorri, Y. Le Cun, and J. Denker · 1992
Earlier work this paper cites.
Probabilistic inference using Markov chain Monte Carlo methods, Sept. 1993
R. M. Neal · 1993
Earlier work this paper cites.
Fast exact multiplication by the Hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Simulating ratios of normalizing constants via a simple identity: A theoretical exploration
X.-L. Meng and W. H. Wong · 1996
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R. M. Neal · 1996
Earlier work this paper cites.
On-line learning in neural networks
L. Bottou · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, and L. K. Saul · 1999
Earlier work this paper cites.
On the convergence of Markovian stochastic algorithms with rapidly decreasing ergodicity rates
L. Younes · 1999
Earlier work this paper cites.
Slice sampling
R. M. Neal · 2000
Earlier work this paper cites.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
Expectation propagation for approximate Bayesian inference
T. P. Minka · 2001
Earlier work this paper cites.
Annealed importance sampling
R. M. Neal · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Information Theory, Inference & Learning Algorithms
D. J. C. MacKay · 2002
Earlier work this paper cites.
Ensemble selection from libraries of models
R. Caruana, A. Niculescu-Mizil, G. Crew, and A. Ksikes · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
A. Hyvärinen · 2005
Earlier work this paper cites.
Estimating ratios of normalizing constants using linked importance sampling
R. M. Neal · 2005
Cited alongside, same era.
Variational message passing
J. Winn and C. M. Bishop · 2005
Cited alongside, same era.
Pattern Recognition and Machine Learning
C. M. Bishop · 2006
Cited alongside, same era.
Model compression
C. Bucilă, R. Caruana, and A. Niculescu-Mizil · 2006
Cited alongside, same era.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Cited alongside, same era.
Some extensions of score matching
A. Hyvärinen · 2007
Cited alongside, same era.
Advances in Markov chain Monte Carlo methods
I. Murray · 2007
Who invented the reverse mode of differentiation?
A. Griewank · 2012
Later among the works it cites.
Deep neural networks for acoustic modeling in speech recognition
G. E. Hinton, L. Deng, D. Yu, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury · 2012
Later among the works it cites.
Machine Learning: A Probabilistic Perspective
K. P. Murphy · 2012
Later among the works it cites.
ADADELTA: an adaptive learning rate method
M. D. Zeiler · 2012
Later among the works it cites.
Speech recognition with deep recurrent neural networks
A. Graves, A.-R. Mohamed, and G. E. Hinton · 2013
Later among the works it cites.
Learning to pass expectation propagation messages
N. Heess, D. Tarlow, and J. Winn · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
EP: A quick reference
T. P. Minka · 2008
Cited alongside, same era.
On the quantitative analysis of Deep Belief Networks
R. Salakhutdinov and I. Murray · 2008
Cited alongside, same era.
Vectorized adaptive quadrature in MATLAB
L. F. Shampine · 2008
Cited alongside, same era.
Graphical models, exponential families, and variational inference
M. J. Wainwright and M. I. Jordan · 2008
Cited alongside, same era.
Learning deep architectures for AI
Y. Bengio · 2009
Cited alongside, same era.
Later among the works it cites.
RNADE: The real-valued neural autoregressive density-estimator
B. Uria, I. Murray, and H. Larochelle · 2013
Later among the works it cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Later among the works it cites.
Findings of the 2014 workshop on statistical machine translation
O. Bojar, C. Buck, C. Federmann, B. Haddow, P. Koehn, J. Leveling, C. Monz, P. Pecina, M. Post, H. Saint-Amand, R. Soricut, L. Specia, and A. Tamchyna · 2014
Later among the works it cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Later among the works it cites.
Learning small-size DNN with output-distribution-based criteria
J. Li, R. Zhao, J.-T. Huang, and Y. Gong · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
A deep and tractable density estimator
B. Uria, I. Murray, and H. Larochelle · 2014
Later among the works it cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Closest in time.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2015
Closest in time.
A. Korattikara, V. Rathod, K. P. Murphy, and M. Welling · 2015
Closest in time.
MNIST handwritten digit database
Y. Le Cun, C. Cortes, and C. J. C. Burges · 2015
Closest in time.
FitNets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Closest in time.
ImageNet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Closest in time.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Closest in time.