Fetching the paper…
Reading the bibliography…
Autoregressive (AR) models have long dominated the landscape of large language models, driving progress across a wide range of tasks.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
M. Roemmele, C. A. Bejan, and A. S. Gordon · 2011
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, Q. N. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
M. G. Johannes Welbl, Nelson F. Liu · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2019
Earlier work this paper cites.
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures
P. J. Ortiz Su’arez, B. Sagot, and L. Romary · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
A survey on image data augmentation for deep learning
C. Shorten and T. M. Khoshgoftaar · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Glu variants improve transformer
N. Shazeer · 2020
Cited alongside, same era.
Structured denoising diffusion models in discrete state-spaces
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg · 2021
Cited alongside, same era.
Autoregressive diffusion models
E. Hoogeboom, A. A. Gritsenko, J. Bastings, B. Poole, R. v. d. Berg, and T. Salimans · 2021
Discrete diffusion modeling by estimating the ratios of the data distribution
A. Lou, C. Meng, and S. Ermon · 2023
Later among the works it cites.
Scaling data-constrained language models
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution, 2024
A. Lou, C. Meng, and S. Ermon · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Cited alongside, same era.
Diffusion-lm improves controllable text generation
X. Li, J. Thickstun, I. Gulrajani, P. S. Liang, and T. B. Hashimoto · 2022
Cited alongside, same era.
Will we run out of data? limits of llm scaling based on human-generated data
P. Villalobos, A. Ho, J. Sevilla, T. Besiroglu, L. Heim, and M. Hobbhahn · 2022
Cited alongside, same era.
Image data augmentation for deep learning: A survey
S. Yang, W. Xiao, M. Zhang, S. Guo, J. Zhao, and F. Shen · 2022
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al · 2023
Cited alongside, same era.
Later among the works it cites.
Scaling up masked diffusion models on text
S. Nie, F. Zhu, C. Du, T. Pang, Q. Liu, G. Zeng, M. Lin, and C. Li · 2024
Later among the works it cites.
σ \sigma -gpts: A new approach to autoregressive models, 2024
A. Pannatier, E. Courdier, and F. Fleuret · 2024
Later among the works it cites.
Simple and effective masked diffusion language models
S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. Chiu, A. Rush, and V. Kuleshov · 2024
Later among the works it cites.
Simplified and generalized masked diffusion for discrete data
J. Shi, K. Han, Z. Wang, A. Doucet, and M. Titsias · 2024
Later among the works it cites.
Industrycorpus2, 2024
X. Shi, L. Zhao, H. Zhou, and D. Hao · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Later among the works it cites.
Randomized autoregressive visual generation
Q. Yu, J. He, X. Deng, X. Shen, and L.-C. Chen · 2024
Later among the works it cites.
Block diffusion: Interpolating between autoregressive and diffusion language models
M. Arriola, A. Gokaslan, J. T. Chiu, Z. Yang, Z. Qi, J. Han, S. S. Sahoo, and V. Kuleshov · 2025
Closest in time.
Large language diffusion models
S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.-R. Wen, and C. Li · 2025
Closest in time.
Unified multimodal discrete diffusion
A. Swerdlow, M. Prabhudesai, S. Gandhi, D. Pathak, and K. Fragkiadaki · 2025
Closest in time.