Fetching the paper…
Reading the bibliography…
Iterative generative models, such as noise conditional score networks and denoising diffusion probabilistic models, produce high quality samples by gradually denoising an initial noise vector.
“Patient Knowledge Distillation for BERT Model Compression”, 2019
Siqi Sun, Yu Cheng, Zhe Gan and Jingjing Liu · 1908
Earlier work this paper cites.
“Well-Read Students Learn Better: On the Importance of Pre-training Compact Models”, 2019
Iulia Turc, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 1908
Earlier work this paper cites.
“DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”, 2020
Victor Sanh, Lysandre Debut, Julien Chaumond and Thomas Wolf · 1910
Earlier work this paper cites.
“Analyzing and Improving the Image Quality of StyleGAN”, 2020
Tero Karras et al · 1912
Earlier work this paper cites.
“Model Compression”
Cristian Bucilua, Rich Caruana and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
“Denoising Diffusion Probabilistic Models”, 2020
Jonathan Ho, Ajay Jain and Pieter Abbeel · 2006
Earlier work this paper cites.
“Training Generative Adversarial Networks with Limited Data”, 2020
Tero Karras et al · 2006
Earlier work this paper cites.
“A tutorial on energy-based learning”
Yann Lecun et al · 2006
Earlier work this paper cites.
“Improved Techniques for Training Score-Based Generative Models”, 2020
Yang Song and Stefano Ermon · 2006
Earlier work this paper cites.
“NVAE: A Deep Hierarchical Variational Autoencoder”, 2020
Arash Vahdat and Jan Kautz · 2007
Earlier work this paper cites.
“Denoising Diffusion Implicit Models”, 2020
Jiaming Song, Chenlin Meng and Stefano Ermon · 2010
Earlier work this paper cites.
“VAEBM: A Symbiosis between Variational Autoencoders and Energy-based Models”, 2020
Zhisheng Xiao, Karsten Kreis, Jan Kautz and Arash Vahdat · 2010
Earlier work this paper cites.
“Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images”, 2020
Rewon Child · 2011
Earlier work this paper cites.
“Score-Based Generative Modeling through Stochastic Differential Equations”, 2020
Yang Song et al · 2011
Earlier work this paper cites.
“A connection between score matching and denoising autoencoders”
Pascal Vincent · 2011
Earlier work this paper cites.
“Learning Energy-Based Models by Diffusion Recovery Likelihood”, 2020
Ruiqi Gao et al · 2012
Earlier work this paper cites.
“Learning Multiple Layers of Features from Tiny Images”
Alex Krizhevsky · 2012
Cited alongside, same era.
“Generative Adversarial Nets”
Ian Goodfellow et al · 2014
Cited alongside, same era.
“Auto-Encoding Variational Bayes”, 2014
Diederik Kingma and Max Welling · 2014
Cited alongside, same era.
“Distilling the Knowledge in a Neural Network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Cited alongside, same era.
“Deep learning face attributes in the wild”
Ziwei Liu, Ping Luo, Xiaogang Wang and Xiaoou Tang · 2015
Cited alongside, same era.
“Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer”
Sergey Zagoruyko and Nikos Komodakis · 2017
Later among the works it cites.
“Born Again Neural Networks”, 2018
Tommaso Furlanello et al · 2018
Later among the works it cites.
“GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”, 2018
Martin Heusel et al · 2018
Later among the works it cites.
“Spectral Normalization for Generative Adversarial Networks”, 2018
Takeru Miyato, Toshiki Kataoka, Masanori Koyama and Yuichi Yoshida · 2018
Later among the works it cites.
“Learning Implicit Generative Models with the Method of Learned Moments”, 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adriana Romero et al · 2015
Cited alongside, same era.
“Deep Unsupervised Learning using Nonequilibrium Thermodynamics”, 2015
Jascha Sohl-Dickstein, Eric. Weiss, Niru Maheswaranathan and Surya Ganguli · 2015
Cited alongside, same era.
“Distilling Knowledge from Ensembles of Neural Networks for Speech Recognition.”
Yevgen Chebotar and Austin Waters · 2016
Cited alongside, same era.
“Improved Techniques for Training GANs”, 2016
Tim Salimans et al · 2016
Cited alongside, same era.
Fisher Yu et al · 2016
Cited alongside, same era.
“Efficient Knowledge Distillation from an Ensemble of Teachers”
T. Fukuda et al · 2017
Cited alongside, same era.
“Adam: A Method for Stochastic Optimization”, 2017
Diederik. Kingma and Jimmy Ba · 2017
Cited alongside, same era.
Suman Ravuri, Shakir Mohamed, Mihaela Rosca and Oriol Vinyals · 2018
Later among the works it cites.
“Large Scale GAN Training for High Fidelity Natural Image Synthesis”
Andrew Brock, Jeff Donahue and Karen Simonyan · 2019
Later among the works it cites.
“Implicit generation and modeling with energy based models”
Yilun Du and Igor Mordatch · 2019
Later among the works it cites.
“Tinybert: Distilling bert for natural language understanding”
Xiaoqi Jiao et al · 2019
Later among the works it cites.
“Knowledge distillation using output errors for self-attention end-to-end models”
Ho-Gyeong Kim et al · 2019
Later among the works it cites.
“Generative modeling by estimating gradients of the data distribution”
Yang Song and Stefano Ermon · 2019
Later among the works it cites.
“Contrastive Representation Distillation”
Yonglong Tian, Dilip Krishnan and Phillip Isola · 2019
Later among the works it cites.
Yan Gao, Titouan Parcollet and Nicholas Lane · 2020
Later among the works it cites.
“Contrastive Distillation on Intermediate Representations for Language Model Compression”
Siqi Sun et al · 2020
Later among the works it cites.
“Improving GAN Training with Probability Ratio Clipping and Sample Reweighting”
Yue Wu et al · 2020
Later among the works it cites.
“Knowledge distillation meets self-supervision”
Guodong Xu, Ziwei Liu, Xiaoxiao Li and Chen Loy · 2020
Later among the works it cites.