Fetching the paper…
Reading the bibliography…
In recent years, particle-based variational inference (ParVI) methods such as Stein variational gradient descent (SVGD) have grown in popularity as scalable methods for Bayesian inference.
On the geometry of Stein variational gradient descent
Duncan, A., Nüsken, N., and Szpruch, L · 1912
Earlier work this paper cites.
A Modern Introduction to Online Learning
Orabona, F · 1912
Earlier work this paper cites.
Opérateurs maximaux monotones et semi-groupes de contractions dans les espaces de Hilbert
Brézis, H · 1973
Earlier work this paper cites.
The performance of universal encoding
Krichevsky, R. E. and Trofimov, V. K · 1981
Earlier work this paper cites.
Minimization Methods for Non-Differentiable Functions
Shor, N. Z · 1985
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Liu, D. C. and Nocedal, J · 1989
Earlier work this paper cites.
Polar factorization and monotone rearrangement of vector-valued functions
Brenier, Y · 1991
Earlier work this paper cites.
On the Convergence of the Proximal Point Algorithm for Convex Minimization
Güler, O · 1991
Earlier work this paper cites.
New problems on minimizing movements
De Giorgi, E · 1993
Earlier work this paper cites.
Independent component analysis, A new concept?
Comon, P · 1994
Earlier work this paper cites.
A New Learning Algorithm for Blind Signal Separation
Amari, S., Cichocki, A., and Yang, H. H · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M · 1996
Earlier work this paper cites.
The Variational Formulation of the Fokker–Planck Equation
Jordan, R., Kinderlehrer, D., and Otto, F · 1998
Earlier work this paper cites.
The Geometry of Dissipative Evolution Equations: The Porous Medium Equation
Otto, F · 2001
Earlier work this paper cites.
An Introduction to MCMC for Machine Learning
Andrieu, C., de Freitas, N., Doucet, A., and Jordan, M. I · 2003
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Slice sampling
Neal, R. M · 2003
Earlier work this paper cites.
Online Convex Programming and Generalized Infinitesimal Gradient Ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Monte Carlo Statistical Methods
Robert, C. P. and Casella, G · 2004
Earlier work this paper cites.
Eulerian Calculus for the Contraction in the Wasserstein Distance
Otto, F. and Westdickenberg, M · 2005
Earlier work this paper cites.
Statistical Mechanics: Algorithms and Computations
Krauth, W · 2006
Earlier work this paper cites.
Gradient Flows: In Metric Spaces and in the Space of Probability Measures
Ambrosio, L., Gigli, N., and Giuseppe Savaré · 2008
Earlier work this paper cites.
Bayesian Probabilistic Matrix Factorization Using Markov Chain Monte Carlo
Salakhutdinov, R. and Mnih, A · 2008
Earlier work this paper cites.
Optimal Transport: Old and New
Villani, C · 2008
Earlier work this paper cites.
Blindness of score-based methods to isolated components and mixing proportions
Wenliang, L. K. and Kanagawa, H · 2008
Earlier work this paper cites.
Monte Carlo Strategies in Scientific Computing
Liu, J. S · 2009
Earlier work this paper cites.
Evolution Equations for Maximal Monotone Operators: Asymptotic Analysis in Continuous and Discrete Time
Peypouquet, J. and Sorin, S · 2010
Earlier work this paper cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
Bauschke, H. H. and Combettes, P. L · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
On the inverse implication of Brenier-Mccann theorems and the structure of (P2(M),W2)
Gigli, N · 2011
Earlier work this paper cites.
Bayesian Learning via Stochastic Gradient Langevin Dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Nonparametric Variational Inference
Gershman, S. J., Hoffman, M. D., and Blei, D. M · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G. E · 2012
Cited alongside, same era.
Bayesian Data Analysis
Gelman, A., Carlin, J. B., Stern, H. S., Dunson, D. B., Vehtari, A., and Rubin, D. B · 2013
Cited alongside, same era.
Energy statistics: A class of statistics based on distances
Székely, G. J. and Rizzo, M. L · 2013
Cited alongside, same era.
Simultaneous Model Selection and Optimization through Parameter-free Stochastic Learning
Orabona, F · 2014
Cited alongside, same era.
The MovieLens Datasets: History and Context
Harper, F. M. and Konstan, J. A · 2015
Cited alongside, same era.
Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
Hernandez-Lobato, J. M. and Adams, R. P · 2015
Cited alongside, same era.
Sampling as optimization in the space of measures: The Langevin dynamics as a composite optimization problem
Wibisono, A · 2018
Later among the works it cites.
Bayesian Model-Agnostic Meta-Learning
Yoon, J., Kim, T., Dia, O., Kim, S., Bengio, Y., and Ahn, S · 2018
Later among the works it cites.
Policy Optimization as Wasserstein Gradient Flows
Zhang, R., Chen, C., Li, C., and Carin, L · 2018
Later among the works it cites.
Bayesian deep convolutional encoder–decoder networks for surrogate modeling and uncertainty quantification
Zhu, Y. and Zabaras, N · 2018
Later among the works it cites.
Message Passing Stein Variational Gradient Descent
Zhuo, J., Liu, C., Shi, J., Zhu, J., Chen, N., and Zhang, B · 2018
Later among the works it cites.
Maximum Mean Discrepancy Gradient Flow
Arbel, M., Korba, A., Salim, A., and Gretton, A · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam: a method for stochastic optimisation
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
A Kernel Test of Goodness of Fit
Chwialkowski, K., Strathmann, H., and Gretton, A · 2016
Cited alongside, same era.
Partial differential equations and stochastic methods in molecular dynamics
Lelièvre, T. and Stoltz, G · 2016
Cited alongside, same era.
Stein Variational Gradient Descent: A General Purpose Bayesian Inference Algorithm
Liu, Q. and Wang, D · 2016
Cited alongside, same era.
A Kernelized Stein Discrepancy for Goodness-of-fit Tests
Liu, Q., Lee, J. D., and Jordan, M · 2016
Cited alongside, same era.
Coin Betting and Parameter-Free Online Learning
Orabona, F. and Pal, D · 2016
Cited alongside, same era.
Later among the works it cites.
User-friendly guarantees for the Langevin Monte Carlo with inaccurate gradient
Dalalyan, A. S. and Karagulyan, A · 2019
Later among the works it cites.
High-dimensional Bayesian inference via the unadjusted Langevin algorithm
Durmus, A. and Moulines, É · 2019
Later among the works it cites.
Analysis of Langevin Monte Carlo via Convex Optimization
Durmus, A., Majewski, S., and Miasojedow, B · 2019
Later among the works it cites.
Parameter-Free Online Convex Optimization with Sub-Exponential Noise
Jun, K.-S. and Orabona, F · 2019
Later among the works it cites.
Understanding and Accelerating Particle-Based Variational Inference
Liu, C., Zhuo, J., Cheng, P., Zhang, R., Zhu, J., and Carin, L · 2019
Later among the works it cites.
Understanding MCMC Dynamics as Flows on the Wasserstein Space
Liu, C., Zhuo, J., and Zhu, J · 2019
Later among the works it cites.
Stein Variational Gradient Descent with Matrix-Valued Kernels
Wang, D., Tang, Z., Bajaj, C., and Liu, Q · 2019
Later among the works it cites.
Projected Stein Variational Gradient Descent
Chen, P. and Ghattas, O · 2020
Later among the works it cites.
SVGD as a kernelized Wasserstein gradient flow of the chi-squared divergence
Chewi, S., Le Gouic, T., Lu, C., Maunu, T., and Rigollet, P · 2020
Later among the works it cites.
A Non-Asymptotic Analysis for Stein Variational Gradient Descent
Korba, A., Salim, A., Arbel, M., Luise, G., and Gretton, A · 2020
Later among the works it cites.
Tutorial on Parameter-Free Online Learning
Orabona, F. and Cutkosky, A · 2020
Later among the works it cites.
The Wasserstein Proximal Gradient Algorithm
Salim, A., Korba, A., and Luise, G · 2020
Later among the works it cites.
Bayesian Deep Learning and a Probabilistic Perspective of Generalization
Wilson, A. G. and Izmailov, P · 2020
Later among the works it cites.
Stein Self-Repulsive Dynamics: Benefits from Past Samples
Ye, M., Ren, T., and Liu, Q · 2020
Later among the works it cites.
Sliced Kernelized Stein Discrepancy
Gong, W., Li, Y., and Hernández-Lobato, J. M · 2021
Later among the works it cites.
Kernel Stein Discrepancy Descent
Korba, A., Pierre-Cyril, Aubin-Frankowski, Majewski, S., and Ablin, P · 2021
Later among the works it cites.
Implicit Parameter-free Online Learning with Truncated Linear Models
Chen, K., Cutkosky, A., and Orabona, F · 2022
Later among the works it cites.
Online Learning to Transport via the Minimal Selection Principle
Guo, W., Hur, Y., Liang, T., and Ryan, C. T · 2022
Later among the works it cites.
Lagrangian Manifold Monte Carlo on Monge Patches
Hartmann, M., Girolami, M., and Klami, A · 2022
Later among the works it cites.
Grassmann Stein Variational Gradient Descent
Liu, X., Zhu, H., Ton, J.-F., Wynne, G., and Duncan, A · 2022
Later among the works it cites.
An n-dimensional Rosenbrock distribution for Markov chain Monte Carlo testing
Pagani, F., Wiegand, M., and Nadarajah, S · 2022
Later among the works it cites.
A Convergence Theory for SVGD in the Population Limit under Talagrand’s Inequality T1
Salim, A., Sun, L., and Peter Richtárik · 2022
Later among the works it cites.
A Finite-Particle Convergence Rate for Stein Variational Gradient Descent
Shi, J. and Mackey, L · 2022
Later among the works it cites.
Improved Stein Variational Gradient Descent with Importance Weights
Sun, L. and Richtárik, P · 2022
Later among the works it cites.
Sampling with Mollified Interaction Energy Descent
Li, L., Liu, Q., Korba, A., Yurochkin, M., and Solomon, J · 2023
Closest in time.