Fetching the paper…
Reading the bibliography…
In recent years, masked diffusion models (MDMs) have emerged as a promising alternative approach for generative modeling over discrete domains.
More on average case vs approximation complexity
Alekhnovich, M · 2003
Earlier work this paper cites.
Estimating random variables from random sparse observations
Montanari, A · 2008
Earlier work this paper cites.
Hiding quiet solutions in random constraint satisfaction problems
Krzakala, F. and Zdeborová, L · 2009
Earlier work this paper cites.
A coupling argument for the random transposition walk
Bormashenko, O · 2011
Earlier work this paper cites.
Asymptotic analysis of the stochastic block model for modular networks and its algorithmic applications
Decelle, A., Krzakala, F., Moore, C., and Zdeborová, L · 2011
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
The benefit of multitask representation learning
Maurer, A., Pontil, M., and Romera-Paredes, B · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
An overview of multi-task learning in deep neural networks
Ruder, S · 2017
Earlier work this paper cites.
Storage capacity in symmetric binary perceptrons
Aubin, B., Perkins, W., and Zdeborová, L · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Probabilistically masked language model capable of autoregressive generation in arbitrary word order
Liao, Y., Jiang, X., and Liu, Q · 2020
Earlier work this paper cites.
3 million sudoku puzzles with ratings, 2020
Radcliffe, D. G · 2020
Earlier work this paper cites.
A reparameterized discrete diffusion model for text generation
Zheng, L., Yuan, J., Yu, L., and Kong, L · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
Austin, J., Johnson, D. D., Ho, J., Tarlow, D., and van den Berg, R · 2021
Earlier work this paper cites.
The overlap gap property: A topological barrier to optimizing over random structures
Gamarnik, D · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Cited alongside, same era.
Provable meta-learning of linear representations
Tripuraneni, N., Jin, C., and Jordan, M. I · 2021
Cited alongside, same era.
Efficient training of language models to fill in the middle, 2022
Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M · 2022
Cited alongside, same era.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
Scaling up masked diffusion models on text
Nie, S., Zhu, F., Du, C., Pang, T., Liu, Q., Zeng, G., Lin, M., and Li, C · 2024
Later among the works it cites.
Your absorbing discrete diffusion secretly models the conditional distributions of clean data
Ou, J., Nie, S., Xue, K., Zhu, F., Sun, J., Li, Z., and Li, C · 2024
Later among the works it cites.
Arrows of time for large language models
Papadopoulos, V., Wenger, J., and Hongler, C · 2024
Later among the works it cites.
Steering masked discrete diffusion models via discrete denoising posterior prediction
Rector-Brooks, J., Hasan, M., Peng, Z., Quinn, Z., Liu, C., Mittal, S., Dziri, N., Bronstein, M., Bengio, Y., Chatterjee, P., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
On statistical inference when fixed points of belief propagation are unstable
Liu, S., Mohanty, S., and Raghavendra, P · 2022
Cited alongside, same era.
Training and inference on any-order autoregressive models the right way
Shih, A., Sadigh, D., and Ermon, S · 2022
Cited alongside, same era.
Slimpajama: A 627b token cleaned and deduplicated version of redpajama, June 2023
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Hardness of sampling solutions from the symmetric binary perceptron
Alaoui, A. E. and Gamarnik, D · 2024
Cited alongside, same era.
Convergence analysis of discrete diffusion model: Exact implementation through uniformization
Chen, H. and Ying, L · 2024
Cited alongside, same era.
Schiff, Y., Sahoo, S. S., Phung, H., Wang, G., Boshar, S., Dalla-torre, H., de Almeida, B. P., Rush, A., Pierrot, T., and Kuleshov, V · 2024
Later among the works it cites.
Causal language modeling can elicit search and reasoning capabilities on logic puzzles
Shah, K., Dikkala, N., Wang, X., and Panigrahy, R · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
Shi, J., Han, K., Wang, Z., Doucet, A., and Titsias, M. K · 2024
Later among the works it cites.
Glauber generative model: Discrete diffusion models via binary classification
Varma, H., Nagaraj, D., and Shanmugam, K · 2024
Later among the works it cites.
Diffusion language models are versatile protein learners
Wang, X., Zheng, Z., Ye, F., Xue, D., Huang, S., and Gu, Q · 2024
Later among the works it cites.
Energy-based diffusion language models for text generation
Xu, M., Geffner, T., Kreis, K., Nie, W., Xu, Y., Leskovec, J., Ermon, S., and Vahdat, A · 2024
Later among the works it cites.
Beyond autoregression: Discrete diffusion for complex reasoning and planning
Ye, J., Gao, J., Gong, S., Zheng, L., Jiang, X., Li, Z., and Kong, L · 2024
Later among the works it cites.
Tinyllama: An open-source small language model
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Later among the works it cites.
Zheng, K., Chen, Y., Mao, H., Liu, M.-Y., Zhu, J., and Zhang, Q · 2024
Later among the works it cites.
The factorization curse: Which tokens you predict underlie the reversal curse and more
Kitouni, O., Nolte, N. S., Williams, A., Rabbat, M., Bouchacourt, D., and Ibrahim, M · 2025
Closest in time.
Large language diffusion models
Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R., and Li, C · 2025
Closest in time.
Path planning for masked diffusion model sampling
Peng, F. Z., Bezemek, Z., Patel, S., Yao, S., Rector-Brooks, J., Tong, A., and Chatterjee, P · 2025
Closest in time.
Simple and effective masked diffusion language models
Sahoo, S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J., Rush, A., and Kuleshov, V · 2025
Closest in time.