Fetching the paper…
Reading the bibliography…
Diffusion models have quickly become the go-to paradigm for generative modelling of perceptual signals (such as images and sound) through iterative refinement.
A theory of the term structure of interest rates
J. C. Cox, J. E. Ingersoll Jr, and S. A. Ross · 1985
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Minimum Bayes-risk decoding for statistical machine translation
S. Kumar and W. Byrne · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
A. Hyvärinen and P. Dayan · 2005
Earlier work this paper cites.
Deep encoder, shallow decoder: Reevaluating non-autoregressive machine translation
J. Kasai, N. Pappas, H. Peng, J. Cross, and N. A. Smith · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
Python 3 Reference Manual
G. Van Rossum and F. L. Drake · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
P. Vincent · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
D. J. Rezende, S. Mohamed, and D. Wierstra · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush · 2016
Earlier work this paper cites.
Non-autoregressive neural machine translation
J. Gu, J. Bradbury, C. Xiong, V. O. Li, and R. Socher · 2017
Earlier work this paper cites.
Using the output embedding to improve language models
O. Press and L. Wolf · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. Van Den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Fast decoding in sequence models using discrete latent variables
L. Kaiser, S. Bengio, A. Roy, A. Vaswani, N. Parmar, J. Uszkoreit, and N. Shazeer · 2018
Earlier work this paper cites.
Glow: Generative flow with invertible 1x1 convolutions
D. P. Kingma and P. Dhariwal · 2018
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
T. Kudo · 2018
Earlier work this paper cites.
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
J. Lee, E. Mansimov, and K. Cho · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
M. Post · 2018
Earlier work this paper cites.
Generalized elbo with constrained optimization, geco
D. J. Rezende and F. Viola · 2018
Earlier work this paper cites.
Parallel wavenet: Fast high-fidelity speech synthesis
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. Driessche, E. Lockhart, L. Cobo, F. Stimberg, et al · 2018
Earlier work this paper cites.
Piano genie
C. Donahue, I. Simon, and S. Dieleman · 2019
Cited alongside, same era.
Neural spline flows
C. Durkan, A. Bekasov, I. Murray, and G. Papamakarios · 2019
Cited alongside, same era.
Mask-predict: Parallel decoding of conditional masked language models
M. Ghazvininejad, O. Levy, Y. Liu, and L. Zettlemoyer · 2019
Cited alongside, same era.
J. Gu, C. Wang, and J. Zhao · 2019
Cited alongside, same era.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Cited alongside, same era.
Neural importance sampling
T. Müller, B. McWilliams, F. Rousselle, M. Gross, and J. Novák · 2019
Cited alongside, same era.
Parallel and flexible sampling from autoregressive models via langevin dynamics
V. Jayaram and J. Thickstun · 2021
Later among the works it cites.
On Neural Differential Equations
P. Kidger · 2021
Later among the works it cites.
Variational diffusion models
D. Kingma, T. Salimans, B. Poole, and J. Ho · 2021
Later among the works it cites.
Improved denoising diffusion probabilistic models
A. Q. Nichol and P. Dhariwal · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
K. Pillutla, S. Swayamdipta, R. Zellers, J. Thickstun, S. Welleck, Y. Choi, and Z. Harchaoui · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Cited alongside, same era.
Insertion transformer: Flexible sequence generation via insertion operations
M. Stern, W. Chan, J. Kiros, and J. Uszkoreit · 2019
Cited alongside, same era.
Bert has a mouth, and it must speak: Bert as a markov random field language model
A. Wang and K. Cho · 2019
Cited alongside, same era.
The DeepMind JAX Ecosystem, 2020
I. Babuschkin, K. Baumli, A. Bell, S. Bhupatiraju, J. Bruce, P. Buchlovsky, D. Budden, T. Cai, A. Clark, I. Danihelka, C. Fantacci, J. Godwin, C. Jones, T. Hennigan, M. Hessel, S. Kapturowski, T. Keck, I. Kemaev, M. King, L. Martens, V. Mikulik, T. Norman, J. Quan, G. Papamakarios, R. Ring, F. Ruiz, A. Sanchez, R. Schneider, E. Sezener, S. Spencer, S. Srinivasan, W. Stokowiec, and F. Viola · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Step-unrolled denoising autoencoders for text generation
N. Savinov, J. Chung, M. Binkowski, E. Elsen, and A. v. d. Oord · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, Y. Lu, S. Pan, B. Wen, and Y. Liu · 2021
Later among the works it cites.
Efficient training of language models to fill in the middle
M. Bavarian, H. Jun, N. Tezak, J. Schulman, C. McLeavey, J. Tworek, and M. Chen · 2022
Closest in time.
AudioLM: a language modeling approach to audio generation
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour · 2022
Closest in time.
A continuous time framework for discrete denoising models
A. Campbell, J. Benton, V. De Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet · 2022
Closest in time.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Closest in time.
Analog bits: Generating discrete data using diffusion models with self-conditioning
T. Chen, R. Zhang, and G. Hinton · 2022
Closest in time.
SSD-LM: Semi-autoregressive simplex-based diffusion language model for text generation and modular control, 2022
X. Han, S. Kumar, and Y. Tsvetkov · 2022
Closest in time.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Closest in time.
Improving non-autoregressive translation models without distillation
X. S. Huang, F. Perez, and M. Volkovs · 2022
Closest in time.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Closest in time.
Diffusion-lm improves controllable text generation
X. L. Li, J. Thickstun, I. Gulrajani, P. Liang, and T. B. Hashimoto · 2022
Closest in time.
Concrete score matching: Generalized score matching for discrete data
C. Meng, K. Choi, J. Song, and S. Ermon · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
Diffuser: Discrete diffusion via edit-based reconstruction, 2022
M. Reid, V. J. Hellendoorn, and G. Neubig · 2022
Closest in time.
Categorical sdes with simplex diffusion
P. H. Richemond, S. Dieleman, and A. Doucet · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Closest in time.
Palette: Image-to-image diffusion models
C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi · 2022
Closest in time.
Training and inference on any-order autoregressive models the right way
A. Shih, D. Sadigh, and S. Ermon · 2022
Closest in time.
Self-conditioned embedding diffusion for text generation, 2022
R. Strudel, C. Tallec, F. Altché, Y. Du, Y. Ganin, A. Mensch, W. Grathwohl, N. Savinov, S. Dieleman, L. Sifre, and R. Leblond · 2022
Closest in time.
Score-based continuous-time discrete diffusion models
H. Sun, L. Yu, B. Dai, D. Schuurmans, and H. Dai · 2022
Closest in time.
Lossless speedup of autoregressive translation with generalized aggressive decoding
H. Xia, T. Ge, F. Wei, and Z. Sui · 2022
Closest in time.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Closest in time.