Fetching the paper…
Reading the bibliography…
Density-ratio estimation via classification is a cornerstone of unsupervised learning.
Neutra-lizing bad geometry in hamiltonian monte carlo using neural transport
Hoffman, M., Sountsov, P., Dillon, J. V., Langmore, I., Tran, D., and Vasudevan, S. (2019) · 1903
Earlier work this paper cites.
Data-efficient image recognition with contrastive predictive coding
Hénaff, O. J., Srinivas, A., De Fauw, J., Razavi, A., Doersch, C., Eslami, S. M., and Oord, A. v. d. (2019) · 1905
Earlier work this paper cites.
Monte Carlo gradient estimation in machine learning
Mohamed, S., Rosca, M., Figurnov, M., and Mnih, A. (2019) · 1906
Earlier work this paper cites.
Fair generative modeling via weak supervision
Grover, A., Choi, K., Shu, R., and Ermon, S. (2019a) · 1910
Earlier work this paper cites.
A Mutual Information Maximization Perspective of Language Representation Learning
Kong, L., d’Autume, C. d. M., Ling, W., Yu, L., Dai, Z., and Yogatama, D. (2019) · 1910
Earlier work this paper cites.
Annealed denoising score matching: Learning energy-based models in high-dimensional spaces
Li, Z., Chen, Y., and Sommer, F. T. (2019) · 1910
Earlier work this paper cites.
Understanding the limitations of variational mutual information estimators
Song, J. and Ermon, S. (2019a) · 1910
Earlier work this paper cites.
Optimization by simulated annealing
Kirkpatrick, S., Gelatt, C., and Vecchi, M. P. (1983) · 1983
Earlier work this paper cites.
Simulating normalizing constants: From importance sampling to bridge sampling to path sampling
Gelman, A. and Meng, X. L. (1998) · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits
LeCun, Y., Cortes, C., and Burges, C. J. (1998) · 1998
Earlier work this paper cites.
Annealed importance sampling
Neal, R. M. (2001) · 2001
Earlier work this paper cites.
On contrastive learning for likelihood-free inference
Durkan, C., Murray, I., and Papamakarios, G. (2020) · 2002
Earlier work this paper cites.
Cutting out the middle-man: Training and evaluating energy-based models without sampling
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., and Zemel, R. (2020) · 2002
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Hinton, G. E. (2002) · 2002
Earlier work this paper cites.
Bayesian experimental design for implicit models by mutual information neural estimation
Kleinegesse, S. and Gutmann, M. U. (2020) · 2002
Earlier work this paper cites.
Energy-Based Processes for Exchangeable Data
Yang, M., Dai, B., Dai, H., and Schuurmans, D. (2020) · 2003
Earlier work this paper cites.
Training deep energy-based models with f-divergence minimization
Yu, L., Song, Y., Song, J., and Ermon, S. (2020) · 2003
Earlier work this paper cites.
Bakhtin, A., Deng, Y., Gross, S., Ott, M., Ranzato, M., and Szlam, A. (2020) · 2004
Earlier work this paper cites.
Learning generative models via discriminative approaches
Tu, Z. (2007) · 2007
Earlier work this paper cites.
Direct importance estimation with model selection and its application to covariate shift adaptation
Sugiyama, M., Nakajima, S., Kashima, H., Buenau, P. V., and Kawanabe, M. (2008) · 2008
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T., Tibshirani, R., and Friedman, J. (2009) · 2009
Earlier work this paper cites.
Deep Boltzmann machines
Salakhutdinov, R. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
Direct density ratio estimation for large-scale covariate shift adaptation
Tsuboi, Y., Kashima, H., Hido, S., Bickel, S., and Sugiyama, M. (2009) · 2009
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
Nguyen, X., Wainwright, M. J., and Jordan, M. I. (2010) · 2010
Earlier work this paper cites.
A family of computationally efficient and simple estimators for unnormalized statistical models
Pihlaja, M., Gutmann, M., and Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Bregman divergence as general framework to estimate unnormalized statistical models
Gutmann, M. and Hirayama, J.-I. (2011) · 2011
Earlier work this paper cites.
Learning deep energy models
Ngiam, J. Z., Chen, P. W. K., and Andrew, Y. N. (2011) · 2011
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Gutmann, M. and Hyvärinen, A. (2012) · 2012
Cited alongside, same era.
Density ratio estimation in machine learning
Sugiyama, M., Suzuki, T., and Kanamori, T. (2012) · 2012
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013) · 2013
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
The No-U-Turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
Hoffman, M. D. and Gelman, A. (2014) · 2014
Conditional Noise-Contrastive Estimation of Unnormalised Models
Ceylan, C. and Gutmann, M. U. (2018) · 2018
Later among the works it cites.
GradNorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Chen, Z., Badrinarayanan, V., Lee, C.-Y., and Rabinovich, A. (2018) · 2018
Later among the works it cites.
Dynamic likelihood-free Inference via Ratio Estimation (DIRE)
Dinev, T. and Gutmann, M. U. (2018) · 2018
Later among the works it cites.
Boosted generative models
Grover, A. and Ermon, S. (2018) · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Kingma, D. P. and Dhariwal, P. (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A note on approximating ABC-MCMC using flexible classifiers
Pham, K. C., Nott, D. J., and Chaudhuri, S. (2014) · 2014
Cited alongside, same era.
Accurate and conservative estimates of MRF log-likelihood using reverse annealing
Burda, Y., Grosse, R., and Salakhutdinov, R. (2015) · 2015
Cited alongside, same era.
A note on the evaluation of generative models
Theis, L., Oord, A. v. d., and Bethge, M. (2015) · 2015
Cited alongside, same era.
A learned representation for artistic style
Dumoulin, V., Shlens, J., and Kudlur, M. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Unsupervised feature extraction by time-contrastive learning and nonlinear ICA
Hyvarinen, A. and Morioka, H. (2016) · 2016
Cited alongside, same era.
Liu, Q., Li, L., Tang, Z., and Zhou, D. (2018) · 2018
Later among the works it cites.
Formal limitations on the measurement of mutual information
McAllester, D. and Stratos, K. (2018) · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y. (2018) · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O. (2018) · 2018
Later among the works it cites.
Noise contrastive estimation: asymptotics, comparison with MC-MLE
Riou-Durand, L. and Chopin, N. (2018) · 2018
Later among the works it cites.
Analysis of noise contrastive estimation from the perspective of asymptotic variance
Uehara, M., Matsuda, T., and Komaki, F. (2018) · 2018
Later among the works it cites.
Learning descriptor networks for 3d shape synthesis and analysis
Xie, J., Zheng, Z., Gao, R., Wang, W., Zhu, S.-C., and Nian Wu, Y. (2018) · 2018
Later among the works it cites.
Exponential family estimation via adversarial dynamics embedding
Dai, B., Liu, Z., Dai, H., He, N., Gretton, A., Song, L., and Schuurmans, D. (2019) · 2019
Later among the works it cites.
Implicit Generation and Modeling with Energy Based Models
Du, Y. and Mordatch, I. (2019) · 2019
Later among the works it cites.
Neural spline flows
Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G. (2019) · 2019
Later among the works it cites.
Your classifier is secretly an energy based model and you should treat it like one
Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K. (2019) · 2019
Later among the works it cites.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. (2019) · 2019
Later among the works it cites.
Nonlinear ICA using auxiliary variables and generalized contrastive learning
Hyvarinen, A., Sasaki, H., and Turner, R. (2019) · 2019
Later among the works it cites.
Efficient Bayesian experimental design for implicit models
Kleinegesse, S. and Gutmann, M. U. (2019) · 2019
Later among the works it cites.
Autoregressive energy machines
Nash, C. and Durkan, C. (2019) · 2019
Later among the works it cites.
Wasserstein dependency measure for representation learning
Ozair, S., Lynch, C., Bengio, Y., Van den Oord, A., Levine, S., and Sermanet, P. (2019) · 2019
Later among the works it cites.
On variational bounds of mutual information
Poole, B., Ozair, S., Van Den Oord, A., Alemi, A., and Tucker, G. (2019) · 2019
Later among the works it cites.
Variational noise-contrastive estimation
Rhodes, B. and Gutmann, M. U. (2019) · 2019
Later among the works it cites.
Self-attention generative adversarial networks
Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A. (2019) · 2019
Later among the works it cites.
Videoflow: A conditional flow-based model for stochastic video generation
Kumar, M., Babaeizadeh, M., Erhan, D., Finn, C., Levine, S., Dinh, L., and Kingma, D. (2020) · 2020
Closest in time.