Fetching the paper…
Reading the bibliography…
In this work, we provide a systematic survey of Discrete Diffusion Language Models (dLLMs) and Discrete Diffusion Multimodal Language Models (dMLLMs).
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in
2015
Earlier work this paper cites.
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep learning with differential privacy,” in
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals
2017
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever
2019
Earlier work this paper cites.
X. Ying, “An overview of overfitting and its solutions,” in
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
L. Floridi and M. Chiriatti, “Gpt-3: Its nature, scope, limits, and consequences,”
2020
Earlier work this paper cites.
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg, “Structured denoising diffusion models in discrete state-spaces,”
2021
Earlier work this paper cites.
E. Hoogeboom, D. Nielsen, P. Jaini, P. Forré, and M. Welling, “Argmax flows and multinomial diffusion: Learning categorical distributions,”
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Campbell, J. Benton, V. De Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet, “A continuous time framework for discrete denoising models,” in
2022
Earlier work this paper cites.
C. Meng, K. Choi, J. Song, and S. Ermon, “Concrete score matching: Generalized score matching for discrete data,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman, “Maskgit: Masked generative image transformer,” in
2022
Earlier work this paper cites.
T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,”
2022
Earlier work this paper cites.
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer, “Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale,”
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray
2022
Earlier work this paper cites.
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,”
2022
Earlier work this paper cites.
C. F. G. D. Santos and J. P. Papa, “Avoiding overfitting: A survey on regularization methods for convolutional neural networks,”
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Sun, L. Yu, B. Dai, D. Schuurmans, and H. Dai, “Score-based continuous-time discrete diffusion models,” in
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
I. Gulrajani and T. B. Hashimoto, “Likelihood-based diffusion language models,”
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023
2023
Earlier work this paper cites.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” 2023
2023
Earlier work this paper cites.
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in
2023
Earlier work this paper cites.
u/emozilla (Reddit user), “Dynamically Scaled RoPE further increases performance of long context LLaMA with zero fine-tuning,” 2023, post on r/LocalLLaMA, Reddit. Available at:
2023
Earlier work this paper cites.
J. X. Morris, W. Zhao, J. T. Chiu, V. Shmatikov, and A. M. Rush, “Language model inversion,”
2023
Earlier work this paper cites.
N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramer, B. Balle, D. Ippolito, and E. Wallace, “Extracting training data from diffusion models,” in
2023
Earlier work this paper cites.
Y. Qu, X. Shen, X. He, M. Backes, S. Zannettou, and Y. Zhang, “Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,” in
2023
Earlier work this paper cites.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang
2023
Earlier work this paper cites.
OpenAI , “Gpt-4o system card,” 2024
2024
Earlier work this paper cites.
OpenAI, “Gpt-4 technical report,” 2024
2024
Earlier work this paper cites.
Gemini Team , “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” 2024
2024
Earlier work this paper cites.
S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. Chiu, A. Rush, and V. Kuleshov, “Simple and effective masked diffusion language models,”
2024
Earlier work this paper cites.
J. Shi, K. Han, Z. Wang, A. Doucet, and M. Titsias, “Simplified and generalized masked diffusion for discrete data,”
2024
Earlier work this paper cites.
S. Gong, S. Agarwal, Y. Zhang, J. Ye, L. Zheng, M. Li, C. An, P. Zhao, W. Bi, J. Han
2024
Earlier work this paper cites.
I. Gat, T. Remez, N. Shaul, F. Kreuk, R. T. Q. Chen, G. Synnaeve, Y. Adi, and Y. Lipman, “Discrete flow matching,” in
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
L. Cheng and S. Li, “Diffuspoll: Conditional text diffusion model for poll generation,” in
2024
Earlier work this paper cites.
Z. Hu, C. Liu, Y. Feng, A. T. Luu, and B. Hooi, “Poetrydiffusion: Towards joint semantic and metrical manipulation in poetry generation,” in
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
L. Zhu, X. Chen, X. Guo, C. Zhang, Z. Zhu, Z. Zhou, and X. Kong, “Pinpointing diffusion grid noise to enhance aspect sentiment quad prediction,” in
2024
Earlier work this paper cites.
S. Iwai, A. Osanai, S. Kitada, and S. Omachi, “Layout-corrector: Alleviating layout sticking phenomenon in discrete diffusion model,” in
2024
Earlier work this paper cites.
J. Ye, S. Gong, L. Chen, L. Zheng, J. Gao, H. Shi, C. Wu, X. Jiang, Z. Li, W. Bi
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
S. Chi, H.-g. Chi, H. Ma, N. Agarwal, F. Siddiqui, K. Ramani, and K. Lee, “M2d2m: Multi-motion generation from text with discrete diffusion models,” in
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
G. Bachmann and V. Nagarajan, “The pitfalls of next-token prediction,” in
P. Huang, S. Liu, Z. Liu, Y. Yan, S. Wang, Z. Chen, and T. Xiao, “Pc-sampler: Position-aware calibration of decoding bias in masked diffusion models,” 2025
2025
Closest in time.
O. Luxembourg, H. Permuter, and E. Nachmani, “Plan for speed: Dilated scheduling for masked diffusion language models,” 2025
2025
Closest in time.
2025
Closest in time.
F. Hong, G. Yu, Y. Ye, H. Huang, H. Zheng, Y. Zhang, Y. Wang, and J. Yao, “Wide-in, narrow-out: Revokable decoding for efficient and effective dllms,” 2025
2025
Closest in time.
X. Ma, R. Yu, G. Fang, and X. Wang, “dkv-cache: The cache for diffusion language models,”
2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu, K. Dang, Y. Fan, Y. Zhang, A. Yang, R. Men, F. Huang, B. Zheng, Y. Miao, S. Quan, Y. Feng, X. Ren, X. Ren, J. Zhou, and J. Lin, “Qwen2.5-coder technical report,” 2024
2024
Cited alongside, same era.
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” January 2024. [Online]. Available:
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
B. F. Labs, “Flux,”
2024
Cited alongside, same era.
J. Bai, T. Ye, W. Chow, E. Song, Q.-G. Chen, X. Li, Z. Dong, L. Zhu, and S. Yan, “Meissonic: Revitalizing masked generative transformers for efficient high-resolution text-to-image synthesis,” in
2024
Cited alongside, same era.
Closest in time.
2025
Closest in time.
C. Wu, H. Zhang, S. Xue, Z. Liu, S. Diao, L. Zhu, P. Luo, S. Han, and E. Xie, “Fast-dllm: Training-free acceleration of diffusion llm by enabling kv cache and parallel decoding,” 2025
2025
Closest in time.
C. Huang and H. Tang, “Ctrldiff: Boosting large diffusion language models with dynamic block prediction and controllable generation,” 2025
2025
Closest in time.
M. Xu, T. Geffner, K. Kreis, W. Nie, Y. Xu, J. Leskovec, S. Ermon, and A. Vahdat, “Energy-based diffusion language models for text generation,” in
2025
Closest in time.
P. Li, Y. Zhou, D. Muhtar, L. Yin, S. Yan, L. Shen, Y. Liang, S. Vosoughi, and S. Liu, “Diffusion language models know the answer before decoding,” 2025
2025
Closest in time.
X. Jin, Y. Wang, Y. Gao, Z. Wen, B. Qi, D. Liu, and L. Zhang, “Thinking inside the mask: In-place prompting in diffusion llms,” 2025
2025
Closest in time.
M. Dang, J. Han, M. Xu, K. Xu, A. Srivastava, and S. Ermon, “Inference-time scaling of diffusion language models with particle gibbs sampling,” 2025
2025
Closest in time.
W. Wang, B. Fang, C. Jing, Y. Shen, Y. Shen, Q. Wang, H. Ouyang, H. Chen, and C. Shen, “Time is a feature: Exploiting temporal dynamics in diffusion language models,” 2025
2025
Closest in time.
X. Liu, Z. Liu, Z. Huang, Q. Guo, Z. He, and X. Qiu, “Longllada: Unlocking long context capabilities in diffusion llms,” 2025
2025
Closest in time.
Y. Song, X. Liu, R. Li, Z. Liu, Z. Huang, Q. Guo, Z. He, and X. Qiu, “Sparse-dllm: Accelerating diffusion llms with dynamic cache eviction,” 2025
2025
Closest in time.
X. Chen, S. Huang, C. Guo, C. Wei, Y. He, J. Zhang, H. H. Li, and Y. Chen, “Dpad: Efficient diffusion language models with suffix dropout,” 2025
2025
Closest in time.
2025
Closest in time.
J. Li, X. Dong, Y. Zang, Y. Cao, J. Wang, and D. Lin, “Beyond fixed: Training-free variable-length denoising for diffusion large language models,” 2025
2025
Closest in time.
2025
Closest in time.
C. Xu and D. Yang, “Dllmquant: Quantizing diffusion-based large language models,”
2025
Closest in time.
Z. Wen, J. Qu, D. Liu, Z. Liu, R. Wu, Y. Yang, X. Jin, H. Xu, X. Liu, W. Li, C. Lu, J. Shao, C. He, and L. Zhang, “The devil behind the mask: An emergent safety vulnerability of diffusion llms,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
D. A. Do, L. A. Tuan, W. Buntine
2025
Closest in time.
W. Shao, M. Liu, and L. Song, “Diffetm: Diffusion process enhanced embedded topic model,”
2025
Closest in time.
X. Dong, W. Li, Y. Le, Z. Jiang, J. Zhong, and Z. Wang, “Termdiffusum: A term-guided diffusion model for extractive summarization of legal documents,” in
2025
Closest in time.
D. Xin, K. Zhao, J. Sun, and Y. Li, “Cdaˆ2: Counterfactual diffusion augmentation for cross-domain adaptation in low-resource sentiment analysis,” in
2025
Closest in time.
Y. Cao, L. Wang, and L. Huang, “Dpcl-diff: Temporal knowledge graph reasoning based on graph node diffusion model with dual-domain periodic contrastive learning,” in
2025
Closest in time.
2025
Closest in time.
E. van Krieken, P. Minervini, E. Ponti, and A. Vergari, “Neurosymbolic diffusion models,”
2025
Closest in time.
2025
Closest in time.
M. Sun, W. Wang, G. Li, J. Liu, J. Sun, W. Feng, S. Lao, S. Zhou, Q. He, and J. Liu, “Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,” in
2025
Closest in time.
A. Jiang, Y. Gao, Z. Sun, Y. Wang, J. Wang, J. Chai, Q. Cao, Y. Heng, H. Jiang, Z. Zhang
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Jung, “Scaffold diffusion: Sparse multi-category voxel structure generation with discrete diffusion,” 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Tang, Y. Zhang, and P. Chatterjee, “Peptune: De novo generation of therapeutic peptides with multi-objective-guided discrete diffusion,”
2025
Closest in time.
2025
Closest in time.
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu, “Qwen2.5 technical report,” 2025
2025
Closest in time.
C. Wu, X. Chen, Z. Wu, Y. Ma, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, C. Ruan
2025
Closest in time.
L. Zhang, “The cosine schedule is fisher-rao-optimal for masked discrete diffusion models,” 2025
2025
Closest in time.
H. He, K. Renz, Y. Cao, and A. Geiger, “Mdpo: Overcoming the training-inference divide of masked diffusion language models,” 2025
2025
Closest in time.
Y. Schiff, S. S. Sahoo, H. Phung, G. Wang, S. Boshar, H. Dalla-torre, B. P. de Almeida, A. M. Rush, T. PIERROT, and V. Kuleshov, “Simple guidance mechanisms for discrete diffusion models,” in
2025
Closest in time.
H. Nisonoff, J. Xiong, S. Allenspach, and J. Listgarten, “Unlocking guidance for discrete state-space diffusion and flow models,” in
2025
Closest in time.
K. Rojas, Y. He, C.-H. Lai, Y. Takida, Y. Mitsufuji, and M. Tao, “Theory-informed improvements to classifier-free guidance for discrete diffusion models,” 2025
2025
Closest in time.
2025
Closest in time.
Y. He, B. Li, L. Liu, Z. Ba, W. Dong, Y. Li, Z. Qin, K. Ren, and C. Chen, “Towards label-only membership inference attack against pre-trained large language models,” in
2025
Closest in time.
2025
Closest in time.