Fetching the paper…
Reading the bibliography…
Discrete diffusion is a promising framework for modeling and generating discrete data.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models
Hyvärinen, A., Hurri, J., Hoyer, P. O., Hyvärinen, A., Hurri, J., and Hoyer, P. O · 2009
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
Nguyen, X., Wainwright, M. J., and Jordan, M. I · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Density-ratio matching under the bregman divergence: a unified framework of density-ratio estimation
Sugiyama, M., Suzuki, T., and Kanamori, T · 2012
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A., Jozefowicz, R., and Bengio, S · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Generative adversarial nets from a density ratio estimation perspective
Uehara, M., Sato, I., Suzuki, M., Nakayama, K., and Matsuo, Y · 2016
Earlier work this paper cites.
Maximum-likelihood augmented discrete generative adversarial networks
Che, T., Li, Y., Zhang, R., Hjelm, R. D., Li, W., Song, Y., and Bengio, Y · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O. K., and Socher, R · 2018
Earlier work this paper cites.
Debiasing evidence approximations: On importance-weighted autoencoders and jackknife variational inference
Nowozin, S · 2018
Earlier work this paper cites.
Training language gans from scratch
de Masson d’Autume, C., Mohamed, S., Rosca, M., and Rae, J. W · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S · 2019
Earlier work this paper cites.
Residual energy-based models for text generation
Deng, Y., Bakhtin, A., Ott, M., Szlam, A., and Ranzato, M · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Learning structured latent factors from dependent data:a generative model framework from information-theoretic perspective
Zhang, R., Koyama, M., and Ishiguro, K · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and van den Berg, R · 2021
Cited alongside, same era.
Argmax flows and multinomial diffusion: Learning categorical distributions
Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., and Welling, M · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
A continuous time framework for discrete denoising models
Campbell, A., Benton, J., Bortoli, V. D., Rainforth, T., Deligiannidis, G., and Doucet, A · 2022
Cited alongside, same era.
Digress: Discrete denoising diffusion for graph generation
Vignac, C., Krawczuk, I., Siraudin, A., Wang, B., Cevher, V., and Frossard, P · 2023
Later among the works it cites.
Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints
Wang, C., Jiang, Y., Yang, C., Liu, H., and Chen, Y · 2023
Later among the works it cites.
Diffusion language models can perform many tasks with scaling and instruction-finetuning
Ye, J., Zheng, Z., Bao, Y., Qian, L., and Gu, Q · 2023
Later among the works it cites.
A reparameterized discrete diffusion model for text generation
Zheng, L., Yuan, J., Yu, L., and Kong, L · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dieleman, S., Sartran, L., Roshannai, A., Savinov, N., Ganin, Y., Richemond, P. H., Doucet, A., Strudel, R., Dyer, C., Durkan, C., Hawthorne, C., Leblond, R., Grathwohl, W., and Adler, J · 2022
Cited alongside, same era.
Diffusion-lm improves controllable text generation
Li, X., Thickstun, J., Gulrajani, I., Liang, P., and Hashimoto, T. B · 2022
Cited alongside, same era.
ParaDetox: Detoxification with parallel data
Logacheva, V., Dementieva, D., Ustyantsev, S., Moskovskiy, D., Dale, D., Krotova, I., Semenov, N., and Panchenko, A · 2022
Cited alongside, same era.
Concrete score matching: Generalized score matching for discrete data
Meng, C., Choi, K., Song, J., and Ermon, S · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
Step-unrolled denoising autoencoders for text generation
Savinov, N., Chung, J., Binkowski, M., Elsen, E., and van den Oord, A · 2022
Cited alongside, same era.
Training and inference on any-order autoregressive models the right way
Shih, A., Sadigh, D., and Ermon, S · 2022
Cited alongside, same era.
Bortoli, V. D., Hutchinson, M. J., Wirnsberger, P., and Doucet, A · 2024
Later among the works it cites.
Campbell, A., Yim, J., Barzilay, R., Rainforth, T., and Jaakkola, T · 2024
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models
Chen, Z., Deng, Y., Yuan, H., Ji, K., and Gu, Q · 2024
Later among the works it cites.
Gat, I., Remez, T., Shaul, N., Kreuk, F., Chen, R. T. Q., Synnaeve, G., Adi, Y., and Lipman, Y · 2024
Later among the works it cites.
Scaling diffusion language models via adaptation from autoregressive models
Gong, S., Agarwal, S., Zhang, Y., Ye, J., Zheng, L., Li, M., An, C., Zhao, P., Bi, W., Han, J., et al · 2024
Later among the works it cites.
Minillm: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2024
Later among the works it cites.
Transfer learning for text diffusion models
Han, K., Kenealy, K., Barua, A., Fiedel, N., and Constant, N · 2024
Later among the works it cites.
Distillm: Towards streamlined distillation for large language models
Ko, J., Kim, S., Chen, T., and Yun, S.-Y · 2024
Later among the works it cites.
Evolving knowledge distillation with large language models and active learning
Liu, C., Zhao, F., Kuang, K., Kang, Y., Jiang, Z., Sun, C., and Wu, F · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution
Lou, A., Meng, C., and Ermon, S · 2024
Later among the works it cites.
Unlocking guidance for discrete state-space diffusion and flow models
Nisonoff, H., Xiong, J., Allenspach, S., and Listgarten, J · 2024
Later among the works it cites.
Your absorbing discrete diffusion secretly models the conditional distributions of clean data
Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z., and Li, C · 2024
Later among the works it cites.
Steering masked discrete diffusion models via discrete denoising posterior prediction
Rector-Brooks, J., Hasan, M., Peng, Z., Quinn, Z., Liu, C., Mittal, S., Dziri, N., Bronstein, M., Bengio, Y., Chatterjee, P., et al · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A. M., and Kuleshov, V · 2024
Later among the works it cites.
Simple guidance mechanisms for discrete diffusion models
Schiff, Y., Sahoo, S. S., Phung, H., Wang, G., Boshar, S., Dalla-torre, H., de Almeida, B. P., Rush, A., Pierrot, T., and Kuleshov, V · 2024
Later among the works it cites.
Flow matching with general discrete paths: A kinetic-optimal perspective
Shaul, N., Gat, I., Havasi, M., Severo, D., Sriram, A., Holderrieth, P., Karrer, B., Lipman, Y., and Chen, R. T · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K · 2024
Later among the works it cites.
Normalizing flows are capable generative models
Zhai, S., Zhang, R., Nakkiran, P., Berthelot, D., Gu, J., Zheng, H., Chen, T., Bautista, M. A., Jaitly, N., and Susskind, J · 2024
Later among the works it cites.
Large language diffusion models
Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R., and Li, C · 2025
Closest in time.
A general framework for inference-time scaling and steering of diffusion models
Singhal, R., Horvitz, Z., Teehan, R., Ren, M., Yu, Z., McKeown, K., and Ranganath, R · 2025
Closest in time.