Fetching the paper…
Reading the bibliography…
We propose a general and scalable approximate sampling strategy for probabilistic models with discrete variables.
Beitrag zur Theorie des Ferround Paramagnetismus
Ising, E · 1924
Earlier work this paper cites.
Statistical analysis of non-lattice data
Besag, J · 1975
Earlier work this paper cites.
Accelerated metropolis method
Umrigar, C. J · 1993
Earlier work this paper cites.
Comments on “representations of knowledge in complex systems” by u. grenander and mi miller
Besag, J · 1994
Earlier work this paper cites.
Peskun’s theorem and a modified discrete-state gibbs sampler
Liu, J. S · 1996
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Gers, F. A., Schmidhuber, J., and Cummins, F · 1999
Earlier work this paper cites.
Correlated mutations in models of protein sequences: phylogenetic and structural effects
Lapedes, A. S., Giraud, B. G., Liu, L., and Stormo, G. D · 1999
Earlier work this paper cites.
Annealed importance sampling
Neal, R. M · 2001
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Hinton, G. E · 2002
Earlier work this paper cites.
The penn treebank: an overview
Taylor, A., Marcus, M., and Santorini, B · 2003
Earlier work this paper cites.
Training restricted boltzmann machines using approximations to the likelihood gradient
Tieleman, T · 2008
Earlier work this paper cites.
Deep belief networks
Hinton, G. E · 2009
Earlier work this paper cites.
Using fast weights to improve persistent contrastive divergence
Tieleman, T. and Hinton, G · 2009
Earlier work this paper cites.
No mcmc for me: Amortized sampling for fast and stable training of energy-based models
Grathwohl, W., Kelly, J., Hashemi, M., Norouzi, M., Swersky, K., and Duvenaud, D · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M. and Hyvärinen, A · 2010
Earlier work this paper cites.
Bayesian models for sparse regression analysis of high dimensional data
Richardson, S., Bottolo, L., and Rosenthal, J. S · 2010
Cited alongside, same era.
Protein 3d structure computed from evolutionary sequence variation
Marks, D. S., Colwell, L. J., Sheridan, R., Hopf, T. A., Pagnani, A., Zecchina, R., and Sander, C · 2011
Cited alongside, same era.
Mcmc using hamiltonian dynamics
Neal, R. M. et al · 2011
Cited alongside, same era.
A kernel two-sample test
Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B., and Smola, A · 2012
Cited alongside, same era.
Interpretation and generalization of score matching
Lyu, S · 2012
Cited alongside, same era.
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Later among the works it cites.
The hamming ball sampler
Titsias, M. K. and Yau, C · 2017
Later among the works it cites.
Stein variational gradient descent without gradient
Han, J. and Liu, Q · 2018
Later among the works it cites.
Learning neural random fields with inclusive auxiliary generators
Song, Y. and Ou, Z · 2018
Later among the works it cites.
Vae with a vampprior
Tomczak, J. and Welling, M · 2018
Later among the works it cites.
Implicit generation and generalization in energy-based models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P. and Welling, M · 2013
Cited alongside, same era.
Adaptive gibbs samplers and related mcmc methods
Łatuszyński, K., Roberts, G. O., Rosenthal, J. S., et al · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Accurate and conservative estimates of mrf log-likelihood using reverse annealing
Burda, Y., Grosse, R., and Salakhutdinov, R · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Liu, Q. and Wang, D · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Cited alongside, same era.
Du, Y. and Mordatch, I · 2019
Later among the works it cites.
Your classifier is secretly an energy based model and you should treat it like one
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K · 2019
Later among the works it cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Later among the works it cites.
Residual energy-based models for text generation
Deng, Y., Bakhtin, A., Ott, M., Szlam, A., and Ranzato, M · 2020
Later among the works it cites.
Stein variational inference for discrete distributions
Han, J., Ding, F., Liu, X., Torresani, L., Peng, J., and Liu, Q · 2020
Later among the works it cites.
Stochastic security: Adversarial defense using long-run dynamics of energy-based models
Hill, M., Mitchell, J., and Zhu, S.-C · 2020
Later among the works it cites.
On the anatomy of mcmc-based maximum likelihood learning of energy-based models
Nijkamp, E., Hill, M., Han, T., Zhu, S.-C., and Wu, Y. N · 2020
Later among the works it cites.
Informed proposals for local mcmc in discrete spaces
Zanella, G · 2020
Later among the works it cites.
Joint energy-based model training for better calibrated natural language understanding models
He, T., McCann, B., Xiong, C., and Hosseini-Asl, E · 2021
Closest in time.