Fetching the paper…
Reading the bibliography…
Encoding domain knowledge into the prior over the high-dimensional weight space of a neural network is challenging but essential in applications with limited data and weak signals.
“Functional variational Bayesian neural networks.”
Sun, S., Zhang, G., Shi, J., and Grosse, R. (2019) · 1903
Earlier work this paper cites.
“Expressive priors in Bayesian neural networks: Kernel combinations and periodic functions.”
Pearce, T., Zaki, M., Brintrup, A., and Neely, A. (2019) · 1905
Earlier work this paper cites.
“Scale mixtures of normal distributions.”
Andrews, D. F. and Mallows, C. L. (1974) · 1974
Earlier work this paper cites.
“Bayesian variable selection in linear regression.”
Mitchell, T. J. and Beauchamp, J. J. (1988) · 1988
Earlier work this paper cites.
Primer of applied regression and analysis of variance
Glantz, S. A., Slinker, B. K., and Neilands, T. B. (1990) · 1990
Earlier work this paper cites.
“A practical Bayesian framework for backpropagation networks.”
MacKay, D. J. (1992) · 1992
Earlier work this paper cites.
“Bayesian non-linear modeling for the prediction competition.”
— (1996) · 1996
Earlier work this paper cites.
“Bayesian model comparison via jump diffusions.”
Phillips, D. B. and Smith, A. F. (1996) · 1996
Earlier work this paper cites.
“Automatic Bayesian curve fitting.”
Denison, D., Mallick, B., and Smith, A. (1998) · 1998
Earlier work this paper cites.
“Feedforward neural networks for nonparametric regression.”
Insua, D. R. and Müller, P. (1998) · 1998
Earlier work this paper cites.
“Robust full Bayesian methods for neural networks.”
Andrieu, C., De Freitas, J. F., and Doucet, A. (2000) · 2000
Earlier work this paper cites.
“On input selection with reversible jump Markov chain Monte Carlo sampling.”
Sykacek, P. (2000) · 2000
Earlier work this paper cites.
“Nonparametric regression using linear combinations of basis functions.”
Kohn, R., Smith, M., and Chan, D. (2001) · 2001
Earlier work this paper cites.
Bayesian model assessment and selection using expected utilities
Vehtari, A. et al. (2001) · 2001
Earlier work this paper cites.
Swiatkowski, J., Roth, K., Veeling, B. S., Tran, L., Dillon, J. V., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S. (2020) · 2002
Earlier work this paper cites.
“Uncertainty Quantification for Sparse Deep Learning.”
Wang, Y. and Ročková, V. (2020) · 2002
Earlier work this paper cites.
“Bayesian deep learning and a probabilistic perspective of generalization.”
Wilson, A. G. and Izmailov, P. (2020) · 2002
Earlier work this paper cites.
“Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors.”
Dusenberry, M. W., Jerfel, G., Wen, Y., Ma, Y.-a., Snoek, J., Heller, K., Lakshminarayanan, B., and Tran, D. (2020) · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. (2006) · 2006
Earlier work this paper cites.
“A general framework for the parametrization of hierarchical models.”
Papaspiliopoulos, O., Roberts, G. O., and Sköld, M. (2007) · 2007
Cited alongside, same era.
“Practical variational inference for neural networks.”
Graves, A. (2011) · 2011
Cited alongside, same era.
“Regression shrinkage and selection via the lasso: a retrospective.”
Tibshirani, R. (2011) · 2011
Cited alongside, same era.
“Imagenet classification with deep convolutional neural networks.”
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Cited alongside, same era.
Bayesian learning for neural networks
Neal, R. M. (2012) · 2012
Cited alongside, same era.
“Reversible jump MCMC simulated annealing for neural networks.”
Andrieu, C., De Freitas, N., and Doucet, A. (2013) · 2013
“Bayesian compression for deep learning.”
Louizos, C., Ullrich, K., and Welling, M. (2017) · 2017
Later among the works it cites.
“Variational dropout sparsifies deep neural networks.”
Molchanov, D., Ashukha, A., and Vetrov, D. (2017) · 2017
Later among the works it cites.
“Structured bayesian pruning via log-normal multiplicative noise.”
Neklyudov, K., Molchanov, D., Ashukha, A., and Vetrov, D. P. (2017) · 2017
Later among the works it cites.
“On the Hyperprior Choice for the Global Shrinkage Parameter in the Horseshoe Prior.”
Piironen, J. and Vehtari, A. (2017) · 2017
Later among the works it cites.
“Sparsity information and regularization in the horseshoe and other shrinkage priors.”
Piironen, J., Vehtari, A., et al. (2017) · 2017
Later among the works it cites.
“Learning structured weight uncertainty in bayesian neural networks.”
Sun, S., Chen, C., and Carin, L. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bayesian Data Analysis
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B. (2013) · 2013
Cited alongside, same era.
“Stochastic variational inference.”
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J. (2013) · 2013
Cited alongside, same era.
“Auto-encoding variational bayes.”
Kingma, D. P. and Welling, M. (2013) · 2013
Cited alongside, same era.
Applied predictive modeling
Kuhn, M., Johnson, K., et al. (2013) · 2013
Cited alongside, same era.
“The horseshoe estimator: Posterior concentration around nearly black vectors.”
Van Der Pas, S. L., Kleijn, B. J., Van Der Vaart, A. W., et al. (2014) · 2014
Cited alongside, same era.
“Hamiltonian Monte Carlo for hierarchical models.”
Betancourt, M. and Girolami, M. (2015) · 2015
Cited alongside, same era.
Later among the works it cites.
“Cohort profile: the National FINRISK study.”
Borodulin, K., Tolonen, H., Jousilahti, P., Jula, A., Juolevi, A., Koskinen, S., Kuulasmaa, K., Laatikainen, T., Männistö, S., Peltonen, M., et al. (2018) · 2018
Later among the works it cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding.”
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Later among the works it cites.
“Structured Variational Learning of Bayesian Neural Networks with Horseshoe Priors.”
Ghosh, S., Yao, J., and Doshi-Velez, F. (2018) · 2018
Later among the works it cites.
“Noise Contrastive Priors for Functional Uncertainty.”
Hafner, D., Tran, D., Lillicrap, T., Irpan, A., and Davidson, J. (2018) · 2018
Later among the works it cites.
“Accurate genomic prediction of human height.”
Lello, L., Avery, S. G., Tellier, L., Vazquez, A. I., de Los Campos, G., and Hsu, S. D. (2018) · 2018
Later among the works it cites.
“Gradient Estimators for Implicit Models.”
Li, Y. and Turner, R. E. (2018) · 2018
Later among the works it cites.
“Posterior concentration for sparse deep learning.”
Polson, N. G. and Ročková, V. (2018) · 2018
Later among the works it cites.
“Variable selection via penalized credible regions with Dirichlet–Laplace global-local shrinkage priors.”
Zhang, Y., Bondell, H. D., et al. (2018) · 2018
Later among the works it cites.
“An Adaptive Empirical Bayesian Method for Sparse Deep Learning.”
Deng, W., Zhang, X., Liang, F., and Lin, G. (2019) · 2019
Later among the works it cites.
“Dropout as a Structured Shrinkage Prior.”
Nalisnick, E., Hernandez-Lobato, J. M., and Smyth, P. (2019) · 2019
Later among the works it cites.
“Bayesian regression using a prior on the model fit: The R2-D2 shrinkage prior.”
Zhang, Y. D., Naughton, B. P., Bondell, H. D., and Reich, B. J. (2020) · 2020
Closest in time.
“Assessing multivariate gene-metabolome associations with rare variants using Bayesian reduced rank regression.”
Marttinen, P., Pirinen, M., Sarin, A.-P., Gillberg, J., Kettunen, J., Surakka, I., Kangas, A. J., Soininen, P., O’Reilly, P., Kaakinen, M., et al. (2014) · 2034
Closest in time.