Fetching the paper…
Reading the bibliography…
Diffusion language models have emerged as a promising approach for text generation.
Perplexity—a measure of the difficulty of speech recognition tasks
Jelinek, F., Mercer, R. L., Bahl, L. R., and Baker, J. K · 1977
Earlier work this paper cites.
Class-based n-gram models of natural language
Brown, P. F., Della Pietra, V. J., Desouza, P. V., Lai, J. C., and Mercer, R. L · 1992
Earlier work this paper cites.
Hidden markov models
Eddy, S. R · 1996
Earlier work this paper cites.
Using a statistical language model to improve the performance of an hmm-based cursive handwriting recognition system
Marti, U.-V. and Bunke, H · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Improved text generation using n-gram statistics
De Novais, E. M., Dias Tadeu, T., and Paraboni, I · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Gpt-3: Its nature, scope, limits, and consequences
Floridi, L. and Chiriatti, M · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
A continuous time framework for discrete denoising models
Campbell, A., Benton, J., Bortoli, V. D., Rainforth, T., Deligiannidis, G., and Doucet, A · 2022
Earlier work this paper cites.
Analog bits: Generating discrete data using diffusion models with self-conditioning
Chen, T., Zhang, R., and Hinton, G · 2022
Earlier work this paper cites.
Continuous diffusion for categorical data
Dieleman, S., Sartran, L., Roshannai, A., Savinov, N., Ganin, Y., Richemond, P. H., Doucet, A., Strudel, R., Dyer, C., Durkan, C., Hawthorne, C., Leblond, R., Grathwohl, W., and Adler, J · 2022
Earlier work this paper cites.
Diffusionbert: Improving generative masked language models with diffusion models
He, Z., Sun, T., Wang, K., Huang, X., and Qiu, X · 2022
Earlier work this paper cites.
On the learning of non-autoregressive transformers
Huang, F., Tao, T., Zhou, H., Li, L., and Huang, M · 2022
Earlier work this paper cites.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Concrete score matching: Generalized score matching for discrete data
Meng, C., Choi, K., Song, J., and Ermon, S · 2022
Cited alongside, same era.
Semi-parametric inducing point networks and neural processes
Rastogi, R., Schiff, Y., Hacohen, A., Li, Z., Lee, I., Deng, Y., Sabuncu, M. R., and Kuleshov, V · 2022
Cited alongside, same era.
Digress: Discrete denoising diffusion for graph generation
Vignac, C., Krawczuk, I., Siraudin, A., Wang, B., Cevher, V., and Frossard, P · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D · 2022
Cited alongside, same era.
Li, Y., Kirchmeyer, A., Mehta, A., Qin, Y., Dadachev, B., Papineni, K., Kumar, S., and Risteski, A · 2024
Later among the works it cites.
Infini-gram: Scaling unbounded n-gram language models to a trillion tokens
Liu, J., Min, S., Zettlemoyer, L., Choi, Y., and Hajishirzi, H · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution
Lou, A., Meng, C., and Ermon, S · 2024
Later among the works it cites.
Diffusion guided language modeling
Lovelace, J., Kishore, V., Chen, Y., and Weinberger, K. Q · 2024
Later among the works it cites.
Beyond perplexity: Examining temporal generalization in large language models via definition generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Dirichlet diffusion score model for biological sequence generation
Avdeyev, P., Shi, C., Tan, Y., Dudnyk, K., and Zhou, J · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Likelihood-based diffusion language models
Gulrajani, I. and Hashimoto, T · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Cited alongside, same era.
Llm is like a box of chocolates: the non-determinism of chatgpt in code generation
Ouyang, S., Zhang, J. M., Harman, M., and Wang, M · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Ellen, X., Adi, Y., Liu, J., Sauvestre, R., Remez, T., et al · 2023
Cited alongside, same era.
Luden, I., Giulianelli, M., and Fernández, R · 2024
Later among the works it cites.
Scaling up masked diffusion models on text
Nie, S., Zhu, F., Du, C., Pang, T., Liu, Q., Zeng, G., Lin, M., and Li, C · 2024
Later among the works it cites.
Your absorbing discrete diffusion secretly models the conditional distributions of clean data
Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z., and Li, C · 2024
Later among the works it cites.
Simple and effective masked diffusion language models, 2024
Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A., and Kuleshov, V · 2024
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Later among the works it cites.
Energy-based diffusion language models for text generation
Xu, M., Geffner, T., Kreis, K., Nie, W., Xu, Y., Leskovec, J., Ermon, S., and Vahdat, A · 2024
Later among the works it cites.
Thought propagation: An analogical approach to complex reasoning with large language models, 2024
Yu, J., He, R., and Ying, R · 2024
Later among the works it cites.
Language rectified flow: Advancing diffusion language generation with probabilistic flows
Zhang, S., Wu, L., Gong, C., and Liu, X · 2024
Later among the works it cites.
Improving and unifying discrete&continuous-time discrete denoising diffusion
Zhao, L., Ding, X., Yu, L., and Akoglu, L · 2024
Later among the works it cites.
Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J., and Zhang, Q · 2024
Later among the works it cites.
Train for the worst, plan for the best: Understanding token ordering in masked diffusions, 2025
Kim, J., Shah, K., Kontonis, V., Kakade, S., and Chen, S · 2025
Closest in time.
Large language diffusion models
Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R., and Li, C · 2025
Closest in time.
Remasking discrete diffusion models with inference-time scaling, 2025
Wang, G., Schiff, Y., Sahoo, S. S., and Kuleshov, V · 2025
Closest in time.
Dream 7b, 2025
Ye, J., Xie, Z., Zheng, L., Gao, J., Wu, Z., Jiang, X., Li, Z., and Kong, L · 2025
Closest in time.
Deductive beam search: Decoding deducible rationale for chain-of-thought reasoning, 2024
Zhu, T., Zhang, K., Xie, J., and Su, Y · 2025
Closest in time.