Fetching the paper…
Reading the bibliography…
Several machine learning applications involve the optimization of higher-order derivatives (e.g., gradients of gradients) during training, which can be expensive in respect to memory and computation even with automatic differentiation.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A scaled conjugate gradient algorithm for fast supervised learning
Martin F Møller · 1990
Earlier work this paper cites.
Some bounds on the complexity of gradients, jacobians, and hessians
Andreas Griewank · 1993
Earlier work this paper cites.
Efficient learning and second-order methods
Yann LeCun · 1993
Earlier work this paper cites.
Mutual information, fisher information, and population coding
Nicolas Brunel and Jean-Pierre Nadal · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
Energy-based models for sparse overcomplete representations
Yee Whye Teh, Max Welling, Simon Osindero, and Geoffrey E Hinton · 2003
Earlier work this paper cites.
Information theory and the central limit theorem
Oliver Johnson · 2004
Earlier work this paper cites.
Analysis 2 springer verlag, 2004
Konrad Königsberger · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Learning nonlinear constraints with contrastive backpropagation
Andriy Mnih and Geoffrey Hinton · 2005
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, M Ranzato, and F Huang · 2006
Earlier work this paper cites.
Uci machine learning repository, 2007
Arthur Asuncion and David Newman · 2007
Earlier work this paper cites.
Evaluating derivatives: principles and techniques of algorithmic differentiation , volume 105
Andreas Griewank and Andrea Walther · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Regularized estimation of image statistics by score matching
Durk P Kingma and Yann L Cun · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng · 2011
Earlier work this paper cites.
Sum-product networks: A new deep architecture
Hoifung Poon and Pedro Domingos · 2011
Earlier work this paper cites.
Wasserstein barycenter and its application to texture mixing
Julien Rabin, Gabriel Peyré, Julie Delon, and Marc Bernot · 2011
Earlier work this paper cites.
Higher order contractive auto-encoder
Salah Rifai, Grégoire Mesnil, Pascal Vincent, Xavier Muller, Yoshua Bengio, Yann Dauphin, and Xavier Glorot · 2011
Earlier work this paper cites.
New method for parameter estimation in probabilistic models: minimum probability flow
Jascha Sohl-Dickstein, Peter B Battaglino, and Michael R DeWeese · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Analysis of numerical methods
Eugene Isaacson and Herbert Bishop Keller · 2012
Earlier work this paper cites.
Estimating the hessian by back-propagating curvature
James Martens, Ilya Sutskever, and Kevin Swersky · 2012
Earlier work this paper cites.
Introduction to numerical analysis , volume 12
Josef Stoer and Roland Bulirsch · 2013
Cited alongside, same era.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Cited alongside, same era.
Clustering via mode seeking by direct estimation of the gradient of a log-density
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Waic, but why? generative ensembles for robust anomaly detection
Hyunsun Choi, Eric Jang, and Alexander A Alemi · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal · 2018
Later among the works it cites.
A large-scale study on regularization and normalization in gans
Karol Kurach, Mario Lucic, Xiaohua Zhai, Marcin Michalski, and Sylvain Gelly · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hiroaki Sasaki, Aapo Hyvärinen, and Masashi Sugiyama · 2014
Cited alongside, same era.
Rotation and scale invariant local binary pattern based on high order directional derivatives for texture classification
Feiniu Yuan · 2014
Cited alongside, same era.
Importance weighted autoencoders
Yuri Burda, Roger Grosse, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Cited alongside, same era.
Gradient-free hamiltonian monte carlo with efficient kernel exponential families
Heiko Strathmann, Dino Sejdinovic, Samuel Livingstone, Zoltan Szabo, and Arthur Gretton · 2015
Cited alongside, same era.
Improving deep neural networks using softplus units
Hao Zheng, Zhanlei Yang, Wenju Liu, Jizhong Liang, and Yanpeng Li · 2015
Cited alongside, same era.
Gradient estimators for implicit models
Yingzhen Li and Richard E Turner · 2018
Later among the works it cites.
Which training methods for gans do actually converge?
Lars Mescheder, Andreas Geiger, and Sebastian Nowozin · 2018
Later among the works it cites.
Deep energy estimator networks
Saeed Saremi, Arash Mehrjou, Bernhard Schölkopf, and Aapo Hyvärinen · 2018
Later among the works it cites.
A spectral approach to gradient estimation for implicit distributions
Jiaxin Shi, Shengyang Sun, and Jun Zhu · 2018
Later among the works it cites.
Efficient and principled score estimation with nyström kernel exponential families
Dougal Sutherland, Heiko Strathmann, Michael Arbel, and Arthur Gretton · 2018
Later among the works it cites.
Implicit generation and modeling with energy based models
Yilun Du and Igor Mordatch · 2019
Later among the works it cites.
Lagging inference networks and posterior collapse in variational autoencoders
Junxian He, Daniel Spokoyny, Graham Neubig, and Taylor Berg-Kirkpatrick · 2019
Later among the works it cites.
Maximum entropy generators for energy-based models
Rithesh Kumar, Sherjil Ozair, Anirudh Goyal, Aaron Courville, and Yoshua Bengio · 2019
Later among the works it cites.
Annealed denoising score matching: Learning energy-based models in high-dimensional spaces
Zengyi Li, Yubei Chen, and Friedrich T Sommer · 2019
Later among the works it cites.
Detecting out-of-distribution inputs to deep generative models using a test for typicality
Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, and Balaji Lakshminarayanan · 2019
Later among the works it cites.
On the anatomy of mcmc-based maximum likelihood learning of energy-based models
Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Neural empirical bayes
Saeed Saremi and Aapo Hyvarinen · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Later among the works it cites.
Sliced score matching: A scalable approach to density and score estimation
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon · 2019
Later among the works it cites.
Learning deep kernels for exponential family densities
Li Wenliang, Dougal Sutherland, Heiko Strathmann, and Arthur Gretton · 2019
Later among the works it cites.
Understanding and stabilizing gans’ training dynamics with control theory
Kun Xu, Chongxuan Li, Huanshu Wei, Jun Zhu, and Bo Zhang · 2019
Later among the works it cites.
Bi-level score matching for learning energy-based latent variable models
Fan Bao, Chongxuan Li, Kun Xu, Hang Su, Jun Zhu, and Bo Zhang · 2020
Closest in time.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Closest in time.
Sumo: Unbiased estimation of log marginal probability for latent variable models
Yucen Luo, Alex Beatson, Mohammad Norouzi, Jun Zhu, David Duvenaud, Ryan P. Adams, and Ricky T. Q. Chen · 2020
Closest in time.
A wasserstein minimum velocity approach to learning unnormalized models
Ziyu Wang, Shuyu Cheng, Yueru Li, Jun Zhu, and Bo Zhang · 2020
Closest in time.
Nonparametric score estimators
Yuhao Zhou, Jiaxin Shi, and Jun Zhu · 2020
Closest in time.