Fetching the paper…
Reading the bibliography…
Many machine learning problems can be expressed as the optimization of some cost functional over a parametric family of probability distributions.
Maximum Mean Discrepancy Gradient Flow
Arbel, M., Korba, A., Salim, A., and Gretton, A. (2019) · 1906
Earlier work this paper cites.
On the mathematical foundations of theoretical statistics
Fisher, R. A. and Russell, E. J. (1922) · 1922
Earlier work this paper cites.
Differential-Geometrical Methods in Statistics
Amari, S.-i. (1985) · 1985
Earlier work this paper cites.
Information and the Accuracy Attainable in the Estimation of Statistical Parameters
Rao, C. R. (1992) · 1992
Earlier work this paper cites.
Natural Gradient Works Efficiently in Learning
Amari, S.-i. (1998) · 1998
Earlier work this paper cites.
Convex Analysis and Variational Problems
Ekeland, I. and Témam, R. (1999) · 1999
Earlier work this paper cites.
A computational fluid mechanics solution to the Monge-Kantorovich mass transfer problem
Benamou, J.-D. and Brenier, Y. (2000) · 2000
Earlier work this paper cites.
On “Natural” Learning and Pruning in Multilayered Perceptrons
Heskes, T. (2000) · 2000
Earlier work this paper cites.
Generalization of an Inequality by Talagrand and Links with the Logarithmic Sobolev Inequality
Otto, F. and Villani, C. (2000) · 2000
Earlier work this paper cites.
The Geometry of Dissipative Evolution Equations: The Porous Medium Equation
Otto, F. (2001) · 2001
Earlier work this paper cites.
A generalized representer theorem
Schölkopf, B., Herbrich, R., and Smola, A. J. (2001) · 2001
Earlier work this paper cites.
Topics in Optimal Transportation
Villani, C. (2003) · 2003
Earlier work this paper cites.
Gradient flows with metric and differentiable structures, and applications to the Wasserstein space
Ambrosio, L., Gigli, N., and Savaré, G. (2004) · 2004
Earlier work this paper cites.
Probability Theory: A Comprehensive Course
Klenke, A. (2008) · 2008
Earlier work this paper cites.
Rodeo: sparse, greedy nonparametric regression
Lafferty, J. and Wasserman, L. (2008) · 2008
Earlier work this paper cites.
Support Vector Machines
Steinwart, I. and Christmann, A. (2008) · 2008
Earlier work this paper cites.
Optimal transport: Old and new
Villani, C. (2009) · 2009
Earlier work this paper cites.
Information geometry of divergence functions
Amari, S. and Cichocki, A. (2010) · 2010
Cited alongside, same era.
Deep Learning via Hessian-free Optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
Adaptive Subgradient Methods for Online Learning and Stochastic Optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Information-geometric optimization algorithms: A unifying picture via invariance principles
Ollivier, Y., Arnold, L., Auger, A., and Hansen, N. (2011) · 2011
Cited alongside, same era.
Lecture 6a overview of mini–batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K. (2012) · 2012
Cited alongside, same era.
Efficient BackProp
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R. (2012) · 2012
Cited alongside, same era.
Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks
Salimans, T. and Kingma, D. P. (2016) · 2016
Later among the works it cites.
Kernel Conditional Exponential Family
Arbel, M. and Gretton, A. (2017) · 2017
Later among the works it cites.
The nonparametric Fisher geometry and the chi-square process density prior
Holbrook, A., Lan, S., Streets, J., and Shahbaba, B. (2017) · 2017
Later among the works it cites.
Geometry of Matrix Decompositions Seen Through Optimal Transport and Information Geometry
Modin, K. (2017) · 2017
Later among the works it cites.
Density estimation in infinite dimensional exponential families
Sriperumbudur, B., Fukumizu, K., Kumar, R., Gretton, A., and Hyvärinen, A. (2017) · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training Deep and Recurrent Networks with Hessian-Free Optimization
Martens, J. and Sutskever, I. (2012) · 2012
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Simonyan, K. and Zisserman, A. (2014) · 2014
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Variational Dropout and the Local Reparameterization Trick
Kingma, D. P., Salimans, T., and Welling, M. (2015) · 2015
Cited alongside, same era.
Later among the works it cites.
Efficient and principled score estimation
Sutherland, D. J., Strathmann, H., Arbel, M., and Gretton, A. (2017) · 2017
Later among the works it cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Bernacchia, A., Lengyel, M., and Hennequin, G. (2018) · 2018
Later among the works it cites.
Natural gradient in Wasserstein statistical manifold
Chen, Y. and Li, W. (2018) · 2018
Later among the works it cites.
Large sample analysis of the median heuristic
Garreau, D., Jitkrittum, W., and Kanagawa, M. (2018) · 2018
Later among the works it cites.
Fast Approximate Natural Gradient Descent in a Kronecker-factored Eigenbasis
George, T., Laurent, C., Bouthillier, X., Ballas, N., and Vincent, P. (2018) · 2018
Later among the works it cites.
Geometry of probability simplex via optimal transport
Li, W. (2018) · 2018
Later among the works it cites.
Wasserstein Riemannian Geometry of Positive Definite Matrices
Malagò, L., Montrucchio, L., and Pistone, G. (2018) · 2018
Later among the works it cites.
SLANG: Fast Structured Covariance Approximations for Bayesian Deep Learning with Natural Gradient
Mishkin, A., Kunstner, F., Nielsen, D., Schmidt, M., and Khan, M. E. (2018) · 2018
Later among the works it cites.
Affine natural proximal learning
Li, W., Lin, A. T., and Montufar, G. (2019) · 2019
Closest in time.
Just Interpolate: Kernel "Ridgeless" Regression Can Generalize
Liang, T. and Rakhlin, A. (2019) · 2019
Closest in time.
Sobolev Descent
Mroueh, Y., Sercu, T., and Raj, A. (2019) · 2019
Closest in time.