Fetching the paper…
Reading the bibliography…
We introduce Autoregressive Diffusion Models (ARDMs), a model class encompassing and generalizing order-agnostic autoregressive models (Uria et al., 2014) and absorbing discrete diffusion (Austin et al., 2021), which we show are special cases of ARDMs under mild assumptions.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 1904
Earlier work this paper cites.
Taking on the curse of dimensionality in joint distributions using neural networks
Samy Bengio and Yoshua Bengio · 2000
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
Jarek Duda · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
The neural autoregressive distribution estimator
Hugo Larochelle and Iain Murray · 2011
Earlier work this paper cites.
Large text compression benchmark, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A deep and tractable density estimator
Benigno Uria, Iain Murray, and Hugo Larochelle · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
FLIF: free lossless image format based on MANIAC compression
Jon Sneyers and Pieter Wuille · 2016
Earlier work this paper cites.
Parallel multiscale autoregressive density estimation
Scott E. Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gomez Colmenarejo, Ziyu Wang, Yutian Chen, Dan Belov, and Nando de Freitas · 2017
Earlier work this paper cites.
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma · 2017
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2018
Earlier work this paper cites.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Constrained graph variational autoencoders for molecule design
Qi Liu, Miltiadis Allamanis, Marc Brockschmidt, and Alexander L. Gaunt · 2018
Earlier work this paper cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2018
Earlier work this paper cites.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Cited alongside, same era.
Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition
P. Warden · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Mask-predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Compression with flows via local bits-back coding
Jonathan Ho, Evan Lohn, and Pieter Abbeel · 2019
Cited alongside, same era.
Integer discrete flows and lossless compression
Predictive sampling with forecasting autoregressive models
Auke J. Wiggers and Emiel Hoogeboom · 2020
Later among the works it cites.
The DEformer: An order-agnostic distribution estimating transformer
Michael A. Alcorn and Anh Nguyen · 2021
Closest in time.
Structured denoising diffusion models in discrete state-spaces
Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg · 2021
Closest in time.
Very deep vaes generalize autoregressive models and can outperform them on images
Rewon Child · 2021
Closest in time.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alex Nichol · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Emiel Hoogeboom, Jorn W. T. Peters, Rianne van den Berg, and Max Welling · 2019
Cited alongside, same era.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling
Jacob Menick and Nal Kalchbrenner · 2019
Cited alongside, same era.
Practical full resolution learned lossless image compression
Fabian Mentzer, Eirikur Agustsson, Michael Tschannen, Radu Timofte, and Luc Van Gool · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Cited alongside, same era.
Practical lossless compression with latent variables using bits back coding
James Townsend, Tom Bird, and David Barber · 2019
Cited alongside, same era.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré, and Max Welling · 2021
Closest in time.
A variational perspective on diffusion-based generative models and score matching
Chin-Wei Huang, Jae Hyun Lim, and Aaron C. Courville · 2021
Closest in time.
Beyond in-place corruption: Insertion and deletion in denoising probabilistic models
Daniel D. Johnson, Jacob Austin, Rianne van den Berg, and Daniel Tarlow · 2021
Closest in time.
Gotta go fast when generating data with score-based models
Alexia Jolicoeur-Martineau, Ke Li, Rémi Piché-Taillefer, Tal Kachman, and Ioannis Mitliagkas · 2021
Closest in time.
Diederik P. Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Closest in time.
On fast sampling of diffusion probabilistic models
Zhifeng Kong and Wei Ping · 2021
Closest in time.
DiffWave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Closest in time.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Closest in time.
Consistency regularization for variational auto-encoders
Samarth Sinha and Adji B. Dieng · 2021
Closest in time.
Accelerating feedforward computation via parallel nonlinear equation solving
Yang Song, Chenlin Meng, Renjie Liao, and Stefano Ermon · 2021
Closest in time.
IDF++: analyzing and improving integer discrete flows for lossless compression
Rianne van den Berg, Alexey A. Gritsenko, Mostafa Dehghani, Casper Kaae Sønderby, and Tim Salimans · 2021
Closest in time.
Learning to efficiently sample from diffusion probabilistic models
Daniel Watson, Jonathan Ho, Mohammad Norouzi, and William Chan · 2021
Closest in time.