Fetching the paper…
Reading the bibliography…
The core principle of Variational Inference (VI) is to convert the statistical inference problem of computing complex posterior probability densities into a tractable optimization problem.
Metric Gaussian variational inference
Knollmüller, J., and Enßlin, T. A. (2019) · 1901
Earlier work this paper cites.
Lee, J., Lee, Y., and Teh, Y. W. (2019b) · 1909
Earlier work this paper cites.
On information and sufficiency
Kullback, S., and Leibler, R. A. (1951) · 1951
Earlier work this paper cites.
A stochastic approximation method
Robbins, H., and Monro, S. (1951) · 1951
Earlier work this paper cites.
Equation of state calculations by fast computing machines
Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H., and Teller, E. (1953) · 1953
Earlier work this paper cites.
A general class of coefficients of divergence of one distribution from another
Ali, S. M., and Silvey, S. D. (1966) · 1966
Earlier work this paper cites.
“Memo” functions and machine learning
Michie, D. (1968) · 1968
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Stein, C. (1972) · 1972
Earlier work this paper cites.
Options: A Monte Carlo approach
Boyle, P. P. (1977) · 1977
Earlier work this paper cites.
Differential geometry of curved exponential families-curvatures and information loss
Amari, S.-I. (1982) · 1982
Earlier work this paper cites.
Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images
Geman, S., and Geman, D. (1984) · 1984
Earlier work this paper cites.
Statistical Field Theory
Parisi, G., and Shankar, R. (1988) · 1988
Earlier work this paper cites.
Approximating probabilistic inference in Bayesian belief networks is NP-hard
Dagum, P., and Luby, M. (1993) · 1993
Earlier work this paper cites.
Riemannian Geometry
do Carmo, M. P. (1993) · 1993
Earlier work this paper cites.
Theoretical Statistics
Cox, D. R., and Hinkley, D. V. (1994) · 1994
Earlier work this paper cites.
Information geometric measurements of generalisation
Zhu, H., and Rohwer, R. (1995) · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M. (1996) · 1996
Earlier work this paper cites.
Mean field theory for sigmoid belief networks
Saul, L. K., Jaakkola, T., and Jordan, M. I. (1996) · 1996
Earlier work this paper cites.
The efficiency and the robustness of natural gradient descent learning rule
Yang, H., and Amari, S.-i. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Tractable variational structures for approximating graphical models
Barber, D., and Wiegerinck, W. (1998) · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Variational probabilistic inference and the QMR-DT network
Jaakkola, T. S., and Jordan, M. I. (1999) · 1999
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., and Jaakkola, Tommi S.and Saul, L. K. (1999) · 1999
Earlier work this paper cites.
LAZY propagation: A junction tree inference algorithm based on lazy evaluation
Madsen, A. L., and Jensen, F. V. (1999) · 1999
Earlier work this paper cites.
Loopy belief propagation for approximate inference: An empirical study
Murphy, K. P., Weiss, Y., and Jordan, M. I. (1999) · 1999
Earlier work this paper cites.
Generalizing variable elimination in Bayesian networks
Gagliardi Cozman, F. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., Mcallester, D., Singh, S., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Mean field methods for a special class of belief networks
Bhattacharyya, C., and Keerthi, S. S. (2001) · 2001
Earlier work this paper cites.
Factor graphs and the sum-product algorithm
Kschischang, F. R., Frey, B. J., and Loeliger, H.-A. (2001) · 2001
Earlier work this paper cites.
Expectation propagation for approximate Bayesian inference
Minka, T. P. (2001) · 2001
Earlier work this paper cites.
Advanced Mean Field Methods: Theory and Practice
Opper, M., and Saad, D. (2001) · 2001
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
MacKay, D. J. C. (2002) · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Schraudolph, N. N. (2002) · 2002
Earlier work this paper cites.
Latent dirichlet allocation
Blei, D. M., Ng, A. Y., and Jordan, M. I. (2003) · 2003
Earlier work this paper cites.
Algorithms for large scale markov blanket discovery
Tsamardinos, I., Aliferis, C. F., and Statnikov, A. R. (2003) · 2003
Earlier work this paper cites.
Convex Optimization
Boyd, S., and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
Graphical models
Jordan, M. I. (2004) · 2004
Earlier work this paper cites.
Divergence measures and message passing
Minka, T. (2005) · 2005
Earlier work this paper cites.
Variational message passing
Winn, J., and Bishop, C. (2005) · 2005
Earlier work this paper cites.
Computing Bayes factors using thermodynamic integration
Lartillot, N., and Philippe, H. (2006) · 2006
Earlier work this paper cites.
Simulation, Fourth Edition
Ross, S. M. (2006) · 2006
Earlier work this paper cites.
Getting started in probabilistic graphical models
Airoldi, E. M. (2007) · 2007
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, M. J., and Jordan, M. I. (2007) · 2007
Earlier work this paper cites.
On some general inequalities related to Jensen’s inequality
Klaričić Bakula, M., Matić, M., and Pečarić, J. (2008) · 2008
Earlier work this paper cites.
The Cramér-Rao inequality
Merberg, A., and Miller, S. J. (2008) · 2008
Earlier work this paper cites.
On the quantitative analysis of deep belief networks
Salakhutdinov, R., and Murray, I. (2008) · 2008
Earlier work this paper cites.
α \alpha -divergence is unique, belonging to both f-divergence and Bregman divergence classes
Amari, S.-I. (2009) · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J. C., Hazan, E., and Singer, Y. (2010) · 2010
Earlier work this paper cites.
Online learning for latent dirichlet allocation
Hoffman, M., Bach, F., and Blei, D. (2010) · 2010
Earlier work this paper cites.
Inductive principles for restricted Boltzmann machine learning
Marlin, B., Swersky, K., Chen, B., and Freitas, N. (2010) · 2010
Earlier work this paper cites.
Density estimation by dual ascent of the log-likelihood
Tabak, E., and Vanden-Eijnden, E. (2010) · 2010
Earlier work this paper cites.
Handbook of Markov Chain Monte Carlo
Brooks, S., Gelman, A., Jones, G., and Meng, X.-L. (2011) · 2011
Earlier work this paper cites.
Probabilistic topic models
Blei, D. M. (2012) · 2012
Earlier work this paper cites.
Nonparametric variational inference
Gershman, S. J., Hoffman, M. D., and Blei, D. M. (2012) · 2012
Cited alongside, same era.
Scalable inference of overlapping communities
Gopalan, P., Mimno, D., Gerrish, S., Freedman, M., and Blei, D. (2012) · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Hinton, G., Srivastava, N., and Swersky, K. (2012) · 2012
Cited alongside, same era.
ADADELTA: An adaptive learning rate method
Zeiler, M. D. (2012) · 2012
Cited alongside, same era.
Stochastic variational inference
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J. (2013) · 2013
Cited alongside, same era.
Machine Learning: A Probabilistic Perspective
Murphy, K. P. (2013) · 2013
Cited alongside, same era.
Improved variational autoencoders for text modeling using dilated convolutions
Yang, Z., Hu, Z., Salakhutdinov, R., and Berg-Kirkpatrick, T. (2017) · 2017
Later among the works it cites.
Quasi-Monte Carlo variational inference
Buchholz, A., Wenzel, F., and Mandt, S. (2018) · 2018
Later among the works it cites.
Understanding disentangling in β \beta -VAE
Burgess, C. P., Higgins, I., Pal, A., Matthey, L., Watters, N., Desjardins, G., and Lerchner, A. (2018) · 2018
Later among the works it cites.
Metrics for deep generative models
Chen, N., Klushyn, A., Kurle, R., Jiang, X., Bayer, J., and Smagt, P. (2018) · 2018
Later among the works it cites.
Inference suboptimality in variational autoencoders
Cremer, C., Li, X., and Duvenaud, D. (2018) · 2018
Later among the works it cites.
Importance sampling for minibatches
Csiba, D., and Richtárik, P. (2018) · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A family of nonparametric density estimation algorithms
Tabak, E. G., and Turner, C. V. (2013) · 2013
Cited alongside, same era.
Variance reduction for stochastic gradient optimization
Wang, C., Chen, X., Smola, A. J., and Xing, E. P. (2013) · 2013
Cited alongside, same era.
Stochastic variational inference for hidden Markov models
Foti, N., Xu, J., Laird, D., and Fox, E. (2014) · 2014
Cited alongside, same era.
Amortized inference in probabilistic reasoning
Gershman, S., and Goodman, N. (2014) · 2014
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Kingma, D. P., and Welling, M. (2014) · 2014
Cited alongside, same era.
Later among the works it cites.
Fast yet simple natural-gradient descent for variational inference in complex models
Khan, M. E., and Nielsen, D. (2018) · 2018
Later among the works it cites.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
Semi-amortized variational autoencoders
Kim, Y., Wiseman, S., Miller, A. C., Sontag, D., and Rush, A. M. (2018) · 2018
Later among the works it cites.
On the challenges of learning with inference networks on sparse, high-dimensional data
Krishnan, R., Liang, D., and Hoffman, M. (2018) · 2018
Later among the works it cites.
Variational autoencoders for collaborative filtering
Liang, D., Krishnan, R. G., Hoffman, M. D., and Jebara, T. (2018) · 2018
Later among the works it cites.
Iterative amortized inference
Marino, J., Yue, Y., and Mandt, S. (2018) · 2018
Later among the works it cites.
Infer.NET 0.3.
Minka, T., Winn, J. M., Guiver, J. P., Zaykov, Y., Fabian, D., and Bronskill, J. (2018) · 2018
Later among the works it cites.
Debiasing evidence approximations: On importance-weighted autoencoders and jackknife variational inference
Nowozin, S. (2018) · 2018
Later among the works it cites.
Tighter variational bounds are not necessarily better
Rainforth, T., Kosiorek, A., Le, T. A., Maddison, C., Igl, M., Wood, F., and Teh, Y. W. (2018) · 2018
Later among the works it cites.
Bayesian convolutional neural networks with variational inference
Shridhar, K., Laumann, F., Maurin, A. L., Olsen, M. A., and Liwicki, M. (2018) · 2018
Later among the works it cites.
Amortized inference regularization
Shu, R., Bui, H. H., Zhao, S., Kochenderfer, M. J., and Ermon, S. (2018) · 2018
Later among the works it cites.
Variational inference: A unified framework of generative models and some revelations
Su, J. (2018) · 2018
Later among the works it cites.
Wasserstein auto-encoders
Tolstikhin, I., Bousquet, O., Gelly, S., and Schoelkopf, B. (2018) · 2018
Later among the works it cites.
Auto-encoding variational neural machine translation
Eikema, B., and Aziz, W. (2019) · 2019
Later among the works it cites.
A style-based generator architecture for generative adversarial networks
Karras, T., Laine, S., and Aila, T. (2019) · 2019
Later among the works it cites.
Understanding posterior collapse in generative latent variable models
Lucas, J., Tucker, G., Grosse, R., and Norouzi, M. (2019) · 2019
Later among the works it cites.
Preventing posterior collapse with delta-VAEs
Razavi, A., van den Oord, A., Poole, B., and Vinyals, O. (2019) · 2019
Later among the works it cites.
Functional variational Bayesian neural networks
Sun, S., Zhang, G., Shi, J., and Grosse, R. (2019) · 2019
Later among the works it cites.
InfoVAE: Balancing learning and inference in variational autoencoders
Zhao, S., Song, J., and Ermon, S. (2019) · 2019
Later among the works it cites.
The autoencoding variational autoencoder
Cemgil, T., Ghaisas, S., Dvijotham, K., Gowal, S., and Kohli, P. (2020) · 2020
Later among the works it cites.
Learning flat latent manifolds with VAEs
Chen, N., Klushyn, A., Ferroni, F., Bayer, J., and Van Der Smagt, P. (2020) · 2020
Later among the works it cites.
Stein variational inference for discrete distributions
Han, J., Ding, F., Liu, X., Torresani, L., Peng, J., and Liu, Q. (2020) · 2020
Later among the works it cites.
Sampling-free variational inference of Bayesian neural networks by variance backpropagation
Haußmann, M., Hamprecht, F. A., and Kandemir, M. (2020) · 2020
Later among the works it cites.
New insights and perspectives on the natural gradient method
Martens, J. (2020) · 2020
Later among the works it cites.
Monte Carlo gradient estimation in machine learning.
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A. (2020) · 2020
Later among the works it cites.
Neural control variates
Müller, T., Rousselle, F., Keller, A., and Novák, J. (2020) · 2020
Later among the works it cites.
Adversarial latent autoencoders
Pidhorskyi, S., Adjeroh, D. A., and Doretto, G. (2020) · 2020
Later among the works it cites.
Dynamics of coordinate ascent variational inference: A case study in 2D Ising models
Plummer, S., Pati, D., and Bhattacharya, A. (2020) · 2020
Later among the works it cites.
Variational learning of Bayesian neural networks via Bayesian dark knowledge
Shen, G., Chen, X., and Deng, Z. (2020) · 2020
Later among the works it cites.
NVAE: A deep hierarchical variational autoencoder
Vahdat, A., and Kautz, J. (2020) · 2020
Later among the works it cites.
A batch normalized inference network keeps the KL vanishing away
Zhu, Q., Bi, W., Liu, X., Ma, X., Li, X., and Wu, D. (2020) · 2020
Later among the works it cites.
Variational inference with Hölder bounds
Chen, J., Lu, D., Xiu, Z., Bai, K., Carin, L., and Tao, C. (2021) · 2021
Later among the works it cites.
An introduction to variational inference
Ganguly, A., and Earp, S. W. (2021) · 2021
Later among the works it cites.
Reducing the amortization gap in variational autoencoders: A Bayesian random function approach
Kim, M., and Pavlovic, V. (2021) · 2021
Later among the works it cites.
Cluster-wise hierarchical generative model for deep amortized clustering
Liu, H., Wang, J., and Jing, L. (2021) · 2021
Later among the works it cites.
Physics enhanced data-driven models with variational Gaussian processes
Marino, D. L., and Manic, M. (2021) · 2021
Later among the works it cites.
Consistency regularization for variational auto-encoders
Sinha, S., and Dieng, A. B. (2021) · 2021
Later among the works it cites.
HypoSVI: Hypocentre inversion with Stein variational inference and physics informed neural networks
Smith, J. D., Ross, Z. E., Azizzadenesheli, K., and Muir, J. B. (2021) · 2021
Later among the works it cites.
Preventing oversmoothing in VAE via generalized variance parameterization
Takida, Y., Liao, W.-H., Lai, C.-H., Uesaka, T., Takahashi, S., and Mitsufuji, Y. (2021) · 2021
Later among the works it cites.
Monte carlo variational auto-encoders
Thin, A., Kotelevskii, N., Doucet, A., Durmus, A., Moulines, E., and Panov, M. (2021) · 2021
Later among the works it cites.
Collapsed variational bounds for Bayesian neural networks
Tomczak, M., Swaroop, S., Foong, A., and Turner, R. (2021) · 2021
Later among the works it cites.
Deep amortized relational model with group-wise hierarchical generative process
Liu, H., Zhou, T., and Wang, J. (2022) · 2022
Closest in time.
Generalization gap in amortized inference
Zhang, M., Hayes, P., and Barber, D. (2022) · 2022
Closest in time.
Tutorial on amortized optimization
Amos, B. (2023) · 2023
Closest in time.
Advances in variational inference
Zhang, C., Butepage, J., Kjellstrom, H., and Mandt, S. (2019) · 2026
Closest in time.
Denoising criterion for variational auto-encoding framework
Im, D. J., Ahn, S., Memisevic, R., and Bengio, Y. (2017) · 2065
Closest in time.
Celeste: Variational inference for a generative model of astronomical images
Regier, J., Miller, A., McAuliffe, J., Adams, R., Hoffman, M., Lang, D., Schlegel, D., and Prabhat, M. (2015) · 2095
Closest in time.