Fetching the paper…
Reading the bibliography…
We propose TraceRL, a trajectory-aware reinforcement learning framework for diffusion language models (DLMs) that incorporates preferred inference trajectory into post-training, and is applicable across different architectures.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Y. Song and S. Ermon · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Structured denoising diffusion models in discrete state-spaces
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman · 2022
Earlier work this paper cites.
Continuous diffusion for categorical data
S. Dieleman, L. Sartran, A. Roshannai, N. Savinov, Y. Ganin, P. H. Richemond, A. Doucet, R. Strudel, C. Dyer, C. Durkan, et al · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. D. Lago, T. Hubert, P. Choy, C. de Masson d’Autume, I. Babuschkin, X. Chen, P.-S. Huang, J. Welbl, S. Gowal, A. Cherepanov, J. Molloy, D. J. Mankowitz, E. S. Robson, P. Kohli, N. de Freitas, K. Kavukcuoglu, and O. Vinyals · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
A. Graves, R. K. Srivastava, T. Atkinson, and F. Gomez · 2023
Earlier work this paper cites.
Likelihood-based diffusion language models
I. Gulrajani and T. B. Hashimoto · 2023
Earlier work this paper cites.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2023
Earlier work this paper cites.
Planner: Generating diversified paragraph via latent language diffusion model
Y. Zhang, J. Gu, Z. Wu, S. Zhai, J. Susskind, and N. Jaitly · 2023
Earlier work this paper cites.
The llama 3 herd of models
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, et al · 2024
Earlier work this paper cites.
Gemini diffusion
Google DeepMind · 2024
Earlier work this paper cites.
V-star: Training verifiers for self-taught reasoners
A. Hosseini, X. Yuan, N. Malkin, A. Courville, A. Sordoni, and R. Agarwal · 2024
Earlier work this paper cites.
S. Jaghouar, J. M. Ong, M. Basra, F. Obeid, J. Straube, M. Keiblinger, E. Bakouch, L. Atkins, M. Panahi, C. Goddard, et al · 2024
Cited alongside, same era.
Livecodebench: Holistic and contamination free evaluation of large language models for code
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica · 2024
Cited alongside, same era.
American invitational mathematics examination (aime) 2024: Aime i and aime ii
Mathematical Association of America, American Mathematics Competitions · 2024
Cited alongside, same era.
Your absorbing discrete diffusion secretly models the conditional distributions of clean data
J. Ou, S. Nie, K. Xue, F. Zhu, J. Sun, Z. Li, and C. Li · 2024
Cited alongside, same era.
Simple and effective masked diffusion language models
Search-r1: Training llms to reason and leverage search engines with reinforcement learning
B. Jin, H. Zeng, Z. Yue, J. Yoon, S. Arik, D. Wang, H. Zamani, and J. Han · 2025
Closest in time.
Train for the worst, plan for the best: Understanding token ordering in masked diffusions
J. Kim, K. Shah, V. Kontonis, S. Kakade, and S. Chen · 2025
Closest in time.
Mercury: Ultra-fast language models based on diffusion
I. Labs, S. Khanna, S. Kharbanda, S. Li, H. Varma, E. Wang, S. Birnbaum, Z. Luo, Y. Miraoui, A. Palrecha, et al · 2025
Closest in time.
dllm-cache: Accelerating diffusion large language models with adaptive caching
Z. Liu, Y. Yang, Y. Zhang, J. Chen, C. Zou, Q. Wei, S. Wang, and L. Zhang · 2025
Closest in time.
Deepcoder: A fully open-source 14b coder at o3-mini level
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. Chiu, A. Rush, and V. Kuleshov · 2024
Cited alongside, same era.
Simplified and generalized masked diffusion for discrete data
J. Shi, K. Han, Z. Wang, A. Doucet, and M. Titsias · 2024
Cited alongside, same era.
Livebench: A challenging, contamination-free llm benchmark
C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Naidu, et al · 2024
Cited alongside, same era.
Beyond autoregression: Discrete diffusion for complex reasoning and planning
J. Ye, J. Gao, S. Gong, L. Zheng, X. Jiang, Z. Li, and L. Kong · 2024
Cited alongside, same era.
Star: Self-taught reasoner bootstrapping reasoning with reasoning
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2024
Cited alongside, same era.
K. Zheng, Y. Chen, H. Mao, M.-Y. Liu, J. Zhu, and Q. Zhang · 2024
Cited alongside, same era.
Block diffusion: Interpolating between autoregressive and diffusion language models
M. Arriola, A. Gokaslan, J. T. Chiu, Z. Yang, Z. Qi, J. Han, S. S. Sahoo, and V. Kuleshov · 2025
Cited alongside, same era.
Sdar: A synergistic diffusion–autoregression paradigm for scalable sequence generation, 2025
S. Cheng, Y. Bian, D. Liu, Y. Jiang, Y. Liu, L. Zhang, W. Wang, Q. Guo, K. Chen, B. Qi*, and B. Zhou · 2025
Cited alongside, same era.
M. Luo, S. Tan, R. Huang, A. Patel, A. Ariyak, Q. Wu, X. Shi, R. Xin, C. Cai, M. Weber, C. Zhang, L. E. Li, R. A. Popa, and I. Stoica · 2025
Closest in time.
dkv-cache: The cache for diffusion language models
X. Ma, R. Yu, G. Fang, and X. Wang · 2025
Closest in time.
Large language diffusion models
S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.-R. Wen, and C. Li · 2025
Closest in time.
Seed diffusion: A large-scale diffusion language model with high-speed inference
Y. Song, Z. Zhang, C. Luo, P. Gao, F. Xia, H. Luo, Z. Li, Y. Yang, H. Yu, X. Qu, et al · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Y. Sui, Y.-N. Chuang, G. Wang, J. Zhang, T. Zhang, J. Yuan, H. Liu, A. Wen, S. Zhong, H. Chen, et al · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Fast-dllm: Training-free acceleration of diffusion llm by enabling kv cache and parallel decoding
C. Wu, H. Zhang, S. Xue, Z. Liu, S. Diao, L. Zhu, P. Luo, S. Han, and E. Xie · 2025
Closest in time.
Dream-coder 7b, 2025
Z. Xie, J. Ye, L. Zheng, J. Gao, J. Dong, Z. Wu, X. Zhao, S. Gong, X. Jiang, Z. Li, and L. Kong · 2025
Closest in time.
Dream 7b, 2025
J. Ye, Z. Xie, L. Zheng, J. Gao, Z. Wu, X. Jiang, Z. Li, and L. Kong · 2025
Closest in time.
d1: Scaling reasoning in diffusion large language models via reinforcement learning
S. Zhao, D. Gupta, Q. Zheng, and A. Grover · 2025
Closest in time.
Llada 1.5: Variance-reduced preference optimization for large language diffusion models
F. Zhu, R. Wang, S. Nie, X. Zhang, C. Wu, J. Hu, J. Zhou, J. Chen, Y. Lin, J.-R. Wen, et al · 2025
Closest in time.
Reasonflux-prm: Trajectory-aware prms for long chain-of-thought reasoning in llms
J. Zou, L. Yang, J. Gu, J. Qiu, K. Shen, J. He, and M. Wang · 2025
Closest in time.