Fetching the paper…
Reading the bibliography…
Energy-Based Models (EBMs), also known as non-normalized probabilistic models, specify probability density or mass functions up to an unknown normalizing constant.
Monte carlo sampling methods using markov chains and their applications
W Keith Hastings · 1970
Earlier work this paper cites.
Correlation functions and computer simulations
Giorgio Parisi · 1981
Earlier work this paper cites.
Estimation of the mean of a multivariate normal distribution
Charles M Stein · 1981
Earlier work this paper cites.
Hybrid monte carlo
Simon Duane, Anthony D Kennedy, Brian J Pendleton, and Duncan Roweth · 1987
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Michael F Hutchinson · 1989
Earlier work this paper cites.
The eigenvalues of mega-dimensional matrices
John Skilling · 1989
Earlier work this paper cites.
Comments on “representations of knowledge in complex systems” by u. grenander and mi miller
JE Besag · 1994
Earlier work this paper cites.
Representations of knowledge in complex systems
Ulf Grenander and Michael I Miller · 1994
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
On the convergence of markovian stochastic algorithms with rapidly decreasing ergodicity rates
Laurent Younes · 1999
Earlier work this paper cites.
Cutting out the middle-man: Training and evaluating energy-based models without sampling
Will Grathwohl, Kuan-Chieh Wang, Jorn-Henrik Jacobsen, David Duvenaud, and Richard Zemel · 2002
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
A new learning algorithm for mean field boltzmann machines
Max Welling and Geoffrey E Hinton · 2002
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Connections between score matching, contrastive divergence, and pseudolikelihood for continuous-valued variables
Aapo Hyvarinen · 2007
Earlier work this paper cites.
Some extensions of score matching
Aapo Hyvärinen · 2007
Earlier work this paper cites.
Learning to be bayesian without supervision
Martin Raphan and Eero P Simoncelli · 2007
Earlier work this paper cites.
A minimum velocity approach to learning
Javier R Movellan · 2008
Earlier work this paper cites.
The graph neural network model
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini · 2008
Earlier work this paper cites.
Training restricted boltzmann machines using approximations to the likelihood gradient
Tijmen Tieleman · 2008
Earlier work this paper cites.
Empirical analysis of the divergence of gibbs sampling based learning algorithms for restricted boltzmann machines
Asja Fischer and Christian Igel · 2010
Earlier work this paper cites.
No mcmc for me: Amortized sampling for fast and stable training of energy-based models
Will Grathwohl, Jacob Kelly, Milad Hashemi, Mohammad Norouzi, Kevin Swersky, and David Duvenaud · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Regularized estimation of image statistics by score matching
Diederik P Kingma and Yann LeCun · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Earlier work this paper cites.
Investigating convergence of restricted boltzmann machine learning
Hannes Schulz, Andreas Müller, and Sven Behnke · 2010
Earlier work this paper cites.
Unifying non-maximum likelihood learning objectives with minimum kl contraction
Siwei Lyu · 2011
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
Learning deep energy models
Jiquan Ngiam, Zhenghao Chen, Pang W Koh, and Andrew Y Ng · 2011
Cited alongside, same era.
Least squares estimation without priors or supervision
Martin Raphan and Eero P Simoncelli · 2011
Cited alongside, same era.
Minimum probability flow learning
Jascha Sohl-Dickstein, Peter Battaglino, and Michael R DeWeese · 2011
Cited alongside, same era.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Cited alongside, same era.
Bregman divergence as general framework to estimate unnormalized statistical models
Michael Gutmann and Jun-ichiro Hirayama · 2012
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Unbiased markov chain monte carlo with couplings
Pierre E Jacob, John O’Leary, and Yves F Atchadé · 2017
Later among the works it cites.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm · 2018
Later among the works it cites.
Adversarial contrastive estimation
Avishek Joey Bose, Huan Ling, and Yanshuai Cao · 2018
Later among the works it cites.
Conditional noise-contrastive estimation of unnormalised models
Ciwan Ceylan and Michael U Gutmann · 2018
Later among the works it cites.
Spherical CNNs
Taco S. Cohen, Mario Geiger, Jonas Köhler, and Max Welling · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Siwei Lyu · 2012
Cited alongside, same era.
Estimating the hessian by back-propagating curvature
James Martens, Ilya Sutskever, and Kevin Swersky · 2012
Cited alongside, same era.
A fast and simple algorithm for training neural probabilistic language models
Andriy Mnih and Yee Whye Teh · 2012
Cited alongside, same era.
Proper local scoring rules
Matthew Parry, A Philip Dawid, Steffen Lauritzen, et al · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Cited alongside, same era.
Monte Carlo theory, methods and examples
Art B. Owen · 2013
Cited alongside, same era.
Later among the works it cites.
Learning generative convnets via multi-grid modeling and sampling
Ruiqi Gao, Yang Lu, Junpei Zhou, Song-Chun Zhu, and Ying Nian Wu · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Cooperative learning of energy-based model and latent variable model via mcmc teaching
Jianwen Xie, Yang Lu, Ruiqi Gao, and Ying Nian Wu · 2018
Later among the works it cites.
Implicit generation and modeling with energy based models
Yilun Du and Igor Mordatch · 2019
Later among the works it cites.
Maximum entropy generators for energy-based models
Rithesh Kumar, Sherjil Ozair, Anirudh Goyal, Aaron Courville, and Yoshua Bengio · 2019
Later among the works it cites.
Learning non-convergent non-persistent short-run mcmc toward energy-based model
Erik Nijkamp, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu · 2019
Later among the works it cites.
Unbiased contrastive divergence algorithm for training energy-based latent variable models
Yixuan Qiu, Lingsong Zhang, and Xiao Wang · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Later among the works it cites.
Sliced score matching: A scalable approach to density and score estimation
Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon · 2019
Later among the works it cites.
Learning deep kernels for exponential family densities
Li Wenliang, Dougal Sutherland, Heiko Strathmann, and Arthur Gretton · 2019
Later among the works it cites.
Learning gradient fields for shape generation
Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge Belongie, Noah Snavely, and Bharath Hariharan · 2020
Later among the works it cites.
Wavegrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan · 2020
Later among the works it cites.
Residual energy-based models for text generation
Yuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio Ranzato · 2020
Later among the works it cites.
Improved contrastive divergence training of energy based models
Yilun Du, Shuang Li, Joshua Tenenbaum, and Igor Mordatch · 2020
Later among the works it cites.
Flow contrastive estimation of energy-based models
Ruiqi Gao, Erik Nijkamp, Diederik P Kingma, Zhen Xu, Andrew M Dai, and Ying Nian Wu · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Later among the works it cites.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2020
Later among the works it cites.
Improved techniques for training score-based generative models
Yang Song and Stefano Ermon · 2020
Later among the works it cites.
A wasserstein minimum velocity approach to learning unnormalized models
Ziyu Wang, Shuyu Cheng, Li Yueru, Jun Zhu, and Bo Zhang · 2020
Later among the works it cites.
Training deep energy-based models with f-divergence minimization
Lantao Yu, Yang Song, Jiaming Song, and Stefano Ermon · 2020
Later among the works it cites.
Score-Based Generative Modeling Through Stochastic Differential Equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Closest in time.