Fetching the paper…
Reading the bibliography…
In arbitrary-order language models, it is an open question how to sample tokens in parallel from the correct joint distribution.
“A mathematical theory of communication”
Claude Shannon · 1948
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“A deep and tractable density estimator”
Benigno Uria, Iain Murray and Hugo Larochelle · 2014
Earlier work this paper cites.
“Made: Masked autoencoder for distribution estimation”
Mathieu Germain, Karol Gregor, Iain Murray and Hugo Larochelle · 2015
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”, 2016
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“A corpus and cloze evaluation for deeper understanding of commonsense stories”
Nasrin Mostafazadeh et al · 2016
Earlier work this paper cites.
“Fixing weight decay regularization in adam”
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Self-attention with relative position representations”
Peter Shaw, Jakob Uszkoreit and Ashish Vaswani · 2018
Earlier work this paper cites.
“OpenWebText Corpus”, http://Skylion007.github.io/OpenWebTextCorpus , 2019
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick and Stefanie Tellex · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Huggingface’s transformers: State-of-the-art natural language processing”
Thomas Wolf et al · 2019
Earlier work this paper cites.
“XLNet: Generalized Autoregressive Pretraining for Language Understanding”
Zhilin Yang · 2019
Earlier work this paper cites.
“Structured denoising diffusion models in discrete state-spaces”
Jacob Austin et al · 2021
Earlier work this paper cites.
“Autoregressive diffusion models”
Emiel Hoogeboom et al · 2021
Earlier work this paper cites.
“Arbitrary conditional distributions with energy”
Ryan Strauss and Junier Oliva · 2021
Earlier work this paper cites.
“Efficient training of language models to fill in the middle”
Mohammad Bavarian et al · 2022
Cited alongside, same era.
“A continuous time framework for discrete denoising models”
Andrew Campbell et al · 2022
Cited alongside, same era.
“Flashattention: Fast and memory-efficient exact attention with io-awareness”
Tri Dao et al · 2022
Cited alongside, same era.
“Incoder: A generative model for code infilling and synthesis”
Daniel Fried et al · 2022
Cited alongside, same era.
“Training and inference on any-order autoregressive models the right way”
Andy Shih, Dorsa Sadigh and Stefano Ermon · 2022
Cited alongside, same era.
“Medusa: Simple llm inference acceleration framework with multiple decoding heads”
Tianle Cai et al · 2024
Later among the works it cites.
“Beyond Autoregression: Fast LLMs via Self-Distillation Through Time”
Justin Deschenaux and Caglar Gulcehre · 2024
Later among the works it cites.
“LayerSkip: Enabling early exit inference and self-speculative decoding”
Mostafa Elhoushi et al · 2024
Later among the works it cites.
“Break the sequential dependency of llm inference using lookahead decoding”
Yichao Fu, Peter Bailis, Ion Stoica and Hao Zhang · 2024
Later among the works it cites.
“Scaling Diffusion Language Models via Adaptation from Autoregressive Models”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam et al · 2023
Cited alongside, same era.
“Accelerating large language model decoding with speculative sampling”
Charlie Chen et al · 2023
Cited alongside, same era.
“Fast inference from transformers via speculative decoding”
Yaniv Leviathan, Matan Kalman and Yossi Matias · 2023
Cited alongside, same era.
“Starcoder: may the source be with you!”
Raymond Li et al · 2023
Cited alongside, same era.
“Discrete diffusion language modeling by estimating the ratios of the data distribution”, 2023
Aaron Lou, Chenlin Meng and Stefano Ermon · 2023
Cited alongside, same era.
“Code llama: Open foundation models for code”
Baptiste Roziere et al · 2023
Cited alongside, same era.
“Gemini: a family of highly capable multimodal models”
Gemini Team et al · 2023
Cited alongside, same era.
Shansan Gong et al · 2024
Later among the works it cites.
“Deepseek-v3 technical report”
Aixin Liu et al · 2024
Later among the works it cites.
“Think While You Generate: Discrete Diffusion with Planned Denoising”
Sulin Liu et al · 2024
Later among the works it cites.
“ σ \sigma -GPTs: A New Approach to Autoregressive Models”
Arnaud Pannatier, Evann Courdier and François Fleuret · 2024
Later among the works it cites.
“Jump your steps: Optimizing sampling schedule of discrete diffusion models”
YH Park et al · 2024
Later among the works it cites.
“Simple and Effective Masked Diffusion Language Models”
Subham Sahoo et al · 2024
Later among the works it cites.
“The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation”
Lawrence Stewart, Matthew Trager, Sujan Gonugondla and Stefano Soatto · 2024
Later among the works it cites.
“Energy-based diffusion language models for text generation”
Minkai Xu et al · 2024
Later among the works it cites.
“Informed correctors for discrete diffusion models”
Yixiu Zhao, Jiaxin Shi, Lester Mackey and Scott Linderman · 2024
Later among the works it cites.
Kaiwen Zheng et al · 2024
Later among the works it cites.
“Large Language Diffusion Models”
Shen Nie et al · 2025
Closest in time.