Fetching the paper…
Reading the bibliography…
In machine learning, it is common to optimize the parameters of a probabilistic model, modulated by an ad hoc regularization term that penalizes some values of the parameters.
Cross-validatory choice and assessment of statistical predictions
Stone, M. (1974) · 1974
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A. (1992) · 1992
Earlier work this paper cites.
Keeping neural networks simple
Hinton, G. E. and van Camp, D. (1993) · 1993
Earlier work this paper cites.
Probable networks and plausible predictions—a review of practical Bayesian methods for supervised neural networks
MacKay, D. J. (1995) · 1995
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by V1?
Olshausen, B. A. and Field, D. J. (1997) · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Handbook of integral equations
Polyanin, A. D. and Manzhirov, A. V. (1998) · 1998
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K. (1999) · 1999
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J. (2003) · 2003
Earlier work this paper cites.
On bayesian classification with Laplace priors
Kaban, A. (2007) · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A. (2011) · 2011
Earlier work this paper cites.
Should penalized least squares regression be interpreted as maximum a posteriori estimation?
Gribonval, R. (2011) · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2013) · 2013
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. L. (2015) · 2015
Cited alongside, same era.
Noisy natural gradient as variational inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R. (2018) · 2018
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
Ashukha, A., Lyzhov, A., Molchanov, D., and Vetrov, D. (2019) · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Later among the works it cites.
Concentration of tempered posteriors and of their variational approximations
Alquier, P., Ridgway, J., et al. (2020) · 2020
Closest in time.
Convergence rates of variational inference in sparse deep learning
Chérief-Abdellatif, B.-E. (2020) · 2020
Closest in time.
Denoising diffusion probabilistic models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P., Salimans, T., and Welling, M. (2015) · 2015
Cited alongside, same era.
Introduction to Fourier analysis on Euclidean spaces
Stein, E. M. and Weiss, G. (2016) · 2016
Cited alongside, same era.
Bayesian compression for deep learning
Louizos, C., Ullrich, K., and Welling, M. (2017) · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D. (2017) · 2017
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S., Gilmer, J., Ganguli, S., and Sohl-Dickstein, J. (2017) · 2017
Cited alongside, same era.
Eigenvalue corrected noisy natural gradient
Bae, J., Zhang, G., and Grosse, R. (2018) · 2018
Cited alongside, same era.
On periodic functions as regularizers for quantization of neural networks
Naumov, M., Diril, U., Park, J., Ray, B., Jablonski, J., and Tulloch, A. (2018) · 2018
Cited alongside, same era.
Ho, J., Jain, A., and Abbeel, P. (2020) · 2020
Closest in time.
How good is the Bayes posterior in deep neural networks really?
Wenzel, F., Roth, K., Veeling, B., Swiatkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S. (2020) · 2020
Closest in time.
Convergence rates of variational posterior distributions
Zhang, F. and Gao, C. (2020) · 2020
Closest in time.
What are bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G. G. (2021) · 2021
Closest in time.
Bayesian neural network priors revisited
Fortuin, V., Garriga-Alonso, A., Ober, S. W., Wenzel, F., Ratsch, G., Turner, R. E., van der Wilk, M., and Aitchison, L. (2022) · 2022
Closest in time.
An optimization-centric view on Bayes’ rule: Reviewing and generalizing variational inference
Knoblauch, J., Jewson, J., and Damoulas, T. (2022) · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Closest in time.