Fetching the paper…
Reading the bibliography…
With the advent of GPU-assisted hardware and maturing high-efficiency software platforms such as TensorFlow and PyTorch, Bayesian posterior sampling for neural networks becomes plausible.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Eigenvalues of covariance matrices: Application to neural-network learning
Y. LeCun, I. Kanter, and S.A. Solla · 1991
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
D.J. MacKay · 1992
Earlier work this paper cites.
Simulated tempering: A new monte carlo scheme
E Marinari and G Parisi · 1992
Earlier work this paper cites.
Bayesian training of backpropagation networks by the hybrid monte carlo method
R.M. Neal · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
G.E. Hinton and D. Van Camp · 1993
Earlier work this paper cites.
Flat Minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Ensemble learning in bayesian neural networks
D. Barber and C. M. Bishop · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Numerical Optimization
J. Nocedal and S.J. Wright · 1999
Earlier work this paper cites.
The concentration of measure phenomenon
M. Ledoux · 2001
Earlier work this paper cites.
Transition path sampling: Throwing ropes over rough mountain passes, in the dark
P.G. Bolhuis, D. Chandler, C. Dellago, and P.L. Geissler · 2002
Earlier work this paper cites.
Large-scale molecular-dynamics simulation of 19 billion particles
K. Kadau, T.C. Germann, and P.S. Lomdahl · 2004
Earlier work this paper cites.
Are loss functions all the same?
Lorenzo Rosasco, Ernesto De Vito, Andrea Caponnetto, Michele Piana, and Alessandro Verri · 2004
Earlier work this paper cites.
Machine learning in computer vision
N. Sebe, I. Cohen, A. Garg, and T.S. Huang · 2005
Earlier work this paper cites.
Diffusion maps
R. Coifman and S. Lafon · 2006
Earlier work this paper cites.
How to generate random matrices from the classical compact groups
F. Mezzadri · 2007
Earlier work this paper cites.
Large-scale molecular dynamics simulations of self-assembling systems
M.L. Klein and W. Shinoda · 2008
Earlier work this paper cites.
Computational complexity of Metropolis-Hastings methods in high dimensions
A. Beskos and A. Stuart · 2009
Earlier work this paper cites.
ACOR package , 2009
J. Goodman · 2009
Earlier work this paper cites.
Long-run accuracy of variational integrators in the stochastic context
N. Bou-Rabee and H. Owhadi · 2010
Earlier work this paper cites.
Ensemble samplers with affine invariance
J. Goodman and J. Weare · 2010
Earlier work this paper cites.
Free Energy Computations
T. Leliévre, M. Rousset, and G. Stoltz · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Bayesian Learning via Stochastic Gradient Langevin Dynamics
M. Welling and Y.-W. Teh · 2011
Cited alongside, same era.
Quasi-stationary distributions: Markov chains, diffusions and dynamical systems
P. Collet, S. Martinez, and J. San Martin · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G.E. Hinton · 2012
Cited alongside, same era.
Rational Construction of Stochastic Numerical Methods for Molecular Sampling
B. Leimkuhler and C. Matthews · 2012
Cited alongside, same era.
Stochastic gradient Riemannian Langevin dynamics on the probability simplex
S. Patterson and Y.-W. Teh · 2013
Cited alongside, same era.
Uncertainty Quantification: Theory, Implementation, and Applications
R. Smith · 2013
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
An empirical analysis of the optimization of deep network loss surfaces
D.J. Im, M. Tao, and K. Branson · 2016
Later among the works it cites.
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir · 2016
Later among the works it cites.
Barzilai-Borwein Step Size for Stochastic Gradient Descent
Conghui Tan, Shiqian Ma, Yu-Hong Dai, and Yuqiu Qian · 2016
Later among the works it cites.
Exploration of the (non-) asymptotic bias and variance of stochastic gradient langevin dynamics
S.J. Vollmer, K.C. Zygalakis, and Y.-W. Teh · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mature hiv-1 capsid structure by cryo-electron microscopy and all-atom molecular dynamics
G. Zhao, J.R. Perilla, E.L. Yufenyuy, X. Meng, B. Chen, J. Ning, J. Ahn, A.M. Greenborn, K. Schulten, C. Aiken, and P. Zhang · 2013
Cited alongside, same era.
Stochastic gradient hamiltonian monte carlo
T. Chen, E. Fox, and C. Guestrin · 2014
Cited alongside, same era.
Qualitatively characterizing neural network optimization problems
I.J. Goodfellow, O. Vinyals, and A.M. Saxe · 2014
Cited alongside, same era.
rmsprop: Divide the gradient by a running average of its recent magnitude, 2014
G. Hinton, N. Srivastava, K. Swervsky, and T. Tieleman · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D.P. Kingma and J. Ba · 2014
Cited alongside, same era.
On the Computational Efficiency of Training Neural Networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Cited alongside, same era.
Later among the works it cites.
Energy landscapes for machine learning
A.J. Ballard, R. Das, S. Martiniani, D. Mehta, L. Sagun, J.D. Stevenson, and D.J. Wales · 2017
Later among the works it cites.
Intrinsic map dynamics exploration for uncharted effective free-energy landscapes
E. Chiavazzoa, R.R. Coifman, R.a Covino, C.W. Gear, A.S. Georgiou, G. Hummer, and I.G. Kevrekidis · 2017
Later among the works it cites.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Later among the works it cites.
Global optimality in neural network training
B.D. Haeffele and R. Vidal · 2017
Later among the works it cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Later among the works it cites.
Generalization in Deep Learning
K. Kawaguchi, L. Kaelbling, and Y. Bengio · 2017
Later among the works it cites.
Recent trends in deep learning based natural language processing
T. Young, D. Hazarika, S. Poria, and E. Cambria · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Later among the works it cites.
Statistics in the big data era: Failures of the machine
D.B. Dunson · 2018
Later among the works it cites.
Artificial intelligence in cardiology
K. Johnson, J. Torres Soto, B. Glicksberg, K. Shameer, R. Miotto, M. Ali, E. Ashley, and J. Dudley · 2018
Later among the works it cites.
Ensemble preconditioning for Markov chain Monte Carlo simulation
B. Leimkuhler, C. Matthews, and J. Weare · 2018
Later among the works it cites.
Visualizing the Loss Landscape of Neural Nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2018
Later among the works it cites.
Langevin Markov Chain Monte Carlo with stochastic gradients
C. Matthews and J. Weare · 2018
Later among the works it cites.
High-dimensional Bayesian inference via the Unadjusted Langevin Algorithm
A. Durmus and E. Moulines · 2019
Closest in time.
Partitioned integrators for thermodynamic parameterization of neural networks
B. Leimkuhler, C. Matthews, and T. Vlaar · 2019
Closest in time.
The simulated tempering method in the infinite switch limit with adaptive weight learning
Anton Martinsson, Jianfeng Lu, Benedict Leimkuhler, and Eric Vanden-Eijnden · 2019
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J.D. Lee · 2019
Closest in time.