Fetching the paper…
Reading the bibliography…
When maximum likelihood estimation is infeasible, one often turns to score matching, contrastive divergence, or minimum probability flow to obtain tractable parameter estimates.
A class of statistics with asymptotically normal distribution
W. Hoeffding · 1948
Earlier work this paper cites.
The strong law of large numbers for U-statistics
W. Hoeffding · 1961
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
C. Stein · 1972
Earlier work this paper cites.
Minimizing a differentiable function over a differential manifold
D. Gabay · 1982
Earlier work this paper cites.
On the convergence of Monte Carlo maximum likelihood calculations
C. J. Geyer · 1994
Earlier work this paper cites.
Large sample estimation and hypothesis testing
W. K. Newey and D. McFadden · 1994
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Muller · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
Sparse code shrinkage: Denoising of nongaussian data by maximum likelihood estimation
A. Hyvärinen · 1999
Earlier work this paper cites.
Statistical Inference
G. Casella and R. Berger · 2001
Earlier work this paper cites.
The Laplace Distribution and Generalizations
S. Kotz, T. J. Kozubowski, and K. Podgorski · 2001
Earlier work this paper cites.
I.-K. Yeo and R. A. Johnson · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
G. E. Hinton · 2002
Earlier work this paper cites.
Learning sparse topographic representations with products of student-t distributions
M. Welling, G. Hinton, and S. Osindero · 2003
Earlier work this paper cites.
Reproducing Kernel Hilbert Spaces in Probability and Statistics
A. Berlinet and C. Thomas-Agnan · 2004
Earlier work this paper cites.
An introduction to Stein’s method
A. Barbour and L. H. Y. Chen · 2005
Earlier work this paper cites.
On learning vector-valued functions
C. A. Micchelli and M. Pontil · 2005
Earlier work this paper cites.
Statistical Inference Based on Divergence Measures , volume 170
L. Pardo · 2005
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
A. Hyvärinen · 2006
Earlier work this paper cites.
Some extensions of score matching
A. Hyvärinen · 2007
Earlier work this paper cites.
Robust Statistics
P. J. Huber and E. M. Ronchetti · 2009
Earlier work this paper cites.
Interpretation and generalization of score matching
S. Lyu · 2009
Earlier work this paper cites.
Fields of experts
S. Roth and M. J. Black · 2009
Earlier work this paper cites.
Minimum probability flow learning
J. Sohl-Dickstein, P. Battaglino, and M. R. DeWeese · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. U. Gutmann and A. Hyvärinen · 2010
Cited alongside, same era.
Regularized estimation of image statistics by score matching
D. P. Kingma and Y. LeCun · 2010
Cited alongside, same era.
A two-layer model of natural stimuli estimated with score matching
U. Köster and A. Hyvärinen · 2010
Cited alongside, same era.
Hilbert space embeddings and metrics on probability measures
B. K. Sriperumbudur, A. Gretton, K. Fukumizu, B. Schölkopf, and G. Lanckriet · 2010
Cited alongside, same era.
Graph kernels
S. V. N. Vishwanathan, N. Schraudolph, R. Kondor, and K. Borgwardt · 2010
Cited alongside, same era.
Statistical Inference: The Minimum Distance Approach
A. Basu, H. Shioya, and C. Park · 2011
Cited alongside, same era.
Stein variational gradient descent: A general purpose Bayesian inference algorithm
Q. Liu and D. Wang · 2016
Later among the works it cites.
A kernelized Stein discrepancy for goodness-of-fit tests
Q. Liu, J. Lee, and M. Jordan · 2016
Later among the works it cites.
Multivariate Stein factors for a class of strongly log-concave distributions
L. Mackey and J. Gorham · 2016
Later among the works it cites.
Score matching estimators for directional distributions
K. V. Mardia, J. T. Kent, and A. K. Laha · 2016
Later among the works it cites.
Operator variational inference
R. Ranganath, J. Altosaar, D. Tran, and D. M. Blei · 2016
Later among the works it cites.
Measuring sample quality with kernels
J. Gorham and L. Mackey · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Normal Approximation by Stein’s Method
L. H. Y. Chen, L. Goldstein, and Q.-M. Shao · 2011
Cited alongside, same era.
Minimum probability flow learning
J. Sohl-dickstein, P. Battaglino, and M. R. DeWeese · 2011
Cited alongside, same era.
On autoencoders and score matching for energy based models
K. Swersky, M. A. Ranzato, D. Buchman, B. M. Marlin, and N. de Freitas · 2011
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
M. U. Gutmann and A. Hyvarinen · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
A. Mnih and Y. W. Teh · 2012
Cited alongside, same era.
Stochastic gradient descent on Riemannian manifolds
S. Bonnabel · 2013
Cited alongside, same era.
Later among the works it cites.
A linear-time kernel goodness-of-fit test
W. Jitkrittum, W. Xu, Z. Szabo, K. Fukumizu, and A. Gretton · 2017
Later among the works it cites.
Riemannian Stein Variational Gradient Descent for Bayesian Inference
C. Liu and J. Zhu · 2017
Later among the works it cites.
Learning deep energy models: Contrastive divergence vs. amortized mle
Q. Liu and D. Wang · 2017
Later among the works it cites.
Black-box Stein divergence minimization for learning latent variable models
C. Ma and D. Barber · 2017
Later among the works it cites.
Control functionals for Monte Carlo integration
C. J. Oates, M. Girolami, and N. Chopin · 2017
Later among the works it cites.
Density estimation in infinite dimensional exponential families
B. Sriperumbudur, K. Fukumizu, A. Gretton, A. Hyvärinen, and R. Kumar · 2017
Later among the works it cites.
Conditional noise-contrastive estimation of unnormalised models
C. Ceylan and M. U. Gutmann · 2018
Later among the works it cites.
Stein points
W. Y. Chen, L. Mackey, J. Gorham, F.-X. Briol, and C. J. Oates · 2018
Later among the works it cites.
A stein variational newton method
G. Detommaso, T. Cui, Y. Marzouk, A. Spantini, and R. Scheichl · 2018
Later among the works it cites.
Learning generative models with Sinkhorn divergences
A. Genevay, G. Peyré, and M. Cuturi · 2018
Later among the works it cites.
Random feature stein discrepancies
J. Huggins and L. Mackey · 2018
Later among the works it cites.
Gradient estimators for implicit models
Y. Li and R. E. Turner · 2018
Later among the works it cites.
Fisher efficient inference of intractable models
S. Liu, T. Kanamori, W. Jitkrittum, and Y. Chen · 2018
Later among the works it cites.
Learning deep kernels for exponential family densities
L. Wenliang, D. Sutherland, H. Strathmann, and A. Gretton · 2018
Later among the works it cites.
Approximate Bayesian computation with the Wasserstein distance
E. Bernton, P. E. Jacob, M. Gerber, and C. P. Robert · 2019
Closest in time.
Statistical inference for generative models with maximum mean discrepancy
F.-X. Briol, A. Barp, A. B. Duncan, and M. Girolami · 2019
Closest in time.
Stein point Markov chain Monte Carlo
W. Y. Chen, A. Barp, F.-X. Briol, J. Gorham, M. Girolami, L. Mackey, and C. J. Oates · 2019
Closest in time.