Fetching the paper…
Reading the bibliography…
Despite their groundbreaking performance for many generative modeling tasks, diffusion models have fallen short on discrete data domains such as natural language.
On a method of investigating periodicities in disturbed series with special reference to wolfer’s sunspot numbers
Yule, G. U · 1971
Earlier work this paper cites.
Reversibility and stochastic networks
Kelly, F · 1980
Earlier work this paper cites.
Stochastic differential equations : an introduction with applications
Øksendal, B · 1987
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Hutchinson, M. F · 1989
Earlier work this paper cites.
Approximate accelerated stochastic simulation of chemically reacting systems
Gillespie, D. T · 2001
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A · 2005
Earlier work this paper cites.
Applied Stochastic Processes and Control for Jump-Diffusions: Modeling, Analysis and Computation
Hanson, F. B · 2007
Earlier work this paper cites.
Some extensions of score matching
Hyvärinen, A · 2007
Earlier work this paper cites.
Tweedie’s formula and selection bias
Efron, B · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Continuous-time Markov chains: An applications-oriented approach
Anderson, W. J · 2012
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Neural networks with cheap differential operators
Chen, R. T. Q. and Duvenaud, D. K · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Cited alongside, same era.
Sliced score matching: A scalable approach to density and score estimation
Song, Y., Garg, S., Shi, J., and Ermon, S · 2019
Cited alongside, same era.
Discrete flows: Invertible generative models of discrete data
Tran, D., Vafa, K., Agrawal, K., Dinh, L., and Poole, B · 2019
Cited alongside, same era.
Bert has a mouth, and it must speak: Bert as a markov random field language model
Wang, A. and Cho, K · 2019
Cited alongside, same era.
Latent normalizing flows for discrete sequences
Ziegler, Z. and Rush, A · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Han, X., Kumar, S., and Tsvetkov, Y · 2022
Later among the works it cites.
Diffusionbert: Improving generative masked language models with diffusion models
He, Z., Sun, T., Wang, K., Huang, X., and Qiu, X · 2022
Later among the works it cites.
Classifier-free diffusion guidance
Ho, J · 2022
Later among the works it cites.
Diffusion-lm improves controllable text generation
Li, X., Thickstun, J., Gulrajani, I., Liang, P. S., and Hashimoto, T. B · 2022
Later among the works it cites.
Concrete score matching: Generalized score matching for discrete data
Meng, C., Choi, K., Song, J., and Ermon, S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A · 2020
Cited alongside, same era.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and van den Berg, R · 2021
Cited alongside, same era.
Argmax flows and multinomial diffusion: Learning categorical distributions
Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., and Welling, M · 2021
Cited alongside, same era.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2021
Cited alongside, same era.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Pillutla, K., Swayamdipta, S., Zellers, R., Thickstun, J., Welleck, S., Choi, Y., and Harchaoui, Z · 2021
Cited alongside, same era.
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
Training and inference on any-order autoregressive models the right way
Shih, A., Sadigh, D., and Ermon, S · 2022
Later among the works it cites.
Self-conditioned embedding diffusion for text generation
Strudel, R., Tallec, C., Altché, F., Du, Y., Ganin, Y., Mensch, A., Grathwohl, W. S., Savinov, N., Dieleman, S., Sifre, L., et al · 2022
Later among the works it cites.
Fast sampling via de-randomization for discrete diffusion models
Chen, Z., Yuan, H., Li, Y., Kou, Y., Zhang, J., and Gu, Q · 2023
Closest in time.
Diffuseq: Sequence to sequence text generation with diffusion models
Gong, S., Li, M., Feng, J., Wu, Z., and Kong, L · 2023
Closest in time.
Graves, A., Srivastava, R. K., Atkinson, T., and Gomez, F · 2023
Closest in time.
Likelihood-based diffusion language models
Gulrajani, I. and Hashimoto, T · 2023
Closest in time.
Reflected diffusion models
Lou, A. and Ermon, S · 2023
Closest in time.
Tess: Text-to-text self-conditioned simplex diffusion
Mahabadi, R. K., Tae, J., Ivison, H., Henderson, J., Beltagy, I., Peters, M. E., and Cohan, A · 2023
Closest in time.
Scalable diffusion models with transformers
Peebles, W. S. and Xie, S · 2023
Closest in time.
Score-based continuous-time discrete diffusion models
Sun, H., Yu, L., Dai, B., Schuurmans, D., and Dai, H · 2023
Closest in time.
Dinoiser: Diffused conditional sequence learning by manipulating noises
Ye, J., Zheng, Z., Bao, Y., Qian, L., and Wang, M · 2023
Closest in time.
A reparameterized discrete diffusion model for text generation
Zheng, L., Yuan, J., Yu, L., and Kong, L · 2023
Closest in time.