Fetching the paper…
Reading the bibliography…
This paper introduces Bayesian Flow Networks (BFNs), a new class of generative model in which the parameters of a set of independent distributions are modified with Bayesian inference in the light of noisy data samples, then passed as input to a neural network that outputs a second, interdependent distribution.
Arithmetic coding for data compression
Ian H Witten, Radford M Neal, and John G Cleary · 1987
Earlier work this paper cites.
Classification by minimum-message-length inference
Chris S. Wallace · 1991
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Conjugate bayesian analysis of the gaussian distribution
Kevin Murphy · 2007
Earlier work this paper cites.
Stochastics: Introduction to Probability and Statistics
H.O. Georgii · 2008
Earlier work this paper cites.
On the quantitative analysis of deep belief networks
Ruslan Salakhutdinov and Iain Murray · 2008
Earlier work this paper cites.
Jarek Duda · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Large text compression benchmark., 2009
Matt Mahoney · 2009
Earlier work this paper cites.
MNIST handwritten digit database, 2010
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P Kingma · 2017
Earlier work this paper cites.
Generating high fidelity images with subscale pixel networks and multidimensional upscaling
Jacob Menick and Nal Kalchbrenner · 2018
Cited alongside, same era.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Pixelvae++: Improved pixelvae with discrete prior
Hossein Sadeghi, Evgeny Andriyash, Walter Vinci, Lorenzo Buffoni, and Mohammad H Amin · 2019
Cited alongside, same era.
Improved denoising diffusion probabilistic models
Alexander Quinn Nichol and Prafulla Dhariwal · 2021
Later among the works it cites.
Consistency regularization for variational auto-encoders
Samarth Sinha and Adji Bousso Dieng · 2021
Later among the works it cites.
Score-based generative modeling in latent space
Arash Vahdat, Karsten Kreis, and Jan Kautz · 2021
Later among the works it cites.
Analog bits: Generating discrete data using diffusion models with self-conditioning
Ting Chen, Ruixiang Zhang, and Geoffrey Hinton · 2022
Later among the works it cites.
Continuous diffusion for categorical data
Sander Dieleman, Laurent Sartran, Arman Roshannai, Nikolay Savinov, Yaroslav Ganin, Pierre H Richemond, Arnaud Doucet, Robin Strudel, Chris Dyer, Conor Durkan, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive Attention Span in Transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Cited alongside, same era.
Practical lossless compression with latent variables using bits back coding
James Townsend, Tom Bird, and David Barber · 2019
Cited alongside, same era.
Discrete flows: Invertible generative models of discrete data
Dustin Tran, Keyon Vafa, Kumar Agrawal, Laurent Dinh, and Ben Poole · 2019
Cited alongside, same era.
Latent Normalizing Flows for Discrete Sequences
Zachary Ziegler and Alexander Rush · 2019
Cited alongside, same era.
Very deep vaes generalize autoregressive models and can outperform them on images
Rewon Child · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Locally masked convolution for autoregressive models
Ajay Jain, Pieter Abbeel, and Deepak Pathak · 2020
Cited alongside, same era.
Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto · 2022
Later among the works it cites.
Categorical SDEs with simplex diffusion
Pierre H. Richemond, Sander Dieleman, and Arnaud Doucet · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho · 2022
Later among the works it cites.
Training and inference on any-order autoregressive models the right way
Andy Shih, Dorsa Sadigh, and Stefano Ermon · 2022
Later among the works it cites.
Self-conditioned embedding diffusion for text generation
Robin Strudel, Corentin Tallec, Florent Altché, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, et al · 2022
Later among the works it cites.
Learning fast samplers for diffusion models by differentiating through sample quality
Daniel Watson, William Chan, Jonathan Ho, and Mohammad Norouzi · 2022
Later among the works it cites.
SSD-LM: Semi-autoregressive simplex-based diffusion language model for text generation and modular control
Xiaochuang Han, Sachin Kumar, and Yulia Tsvetkov · 2023
Closest in time.
Aaron Lou and Stefano Ermon · 2023
Closest in time.
Tess: Text-to-text self-conditioned simplex diffusion
Rabeeh Karimi Mahabadi, Jaesung Tae, Hamish Ivison, James Henderson, Iz Beltagy, Matthew E. Peters, and Arman Cohan · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever · 2023
Closest in time.