Fetching the paper…
Reading the bibliography…
The structure of causal language model training assumes that each token can be accurately predicted from the previous context.
Depth-first search and linear graph algorithms
Robert Tarjan · 1972
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J Williams and David Zipser · 1989
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Anticipating visual representations from unlabeled video
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher · 2017
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2019
Earlier work this paper cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Earlier work this paper cites.
Ctrl: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Non-monotonic sequential text generation
Sean Welleck, Kianté Brantley, Hal Daumé Iii, and Kyunghyun Cho · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Gedi: Generative discriminator guided sequence generation
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani · 2020
Earlier work this paper cites.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Cited alongside, same era.
Kushal Arora, Layla El Asri, Hareesh Bahuleyan, and Jackie Chi Kit Cheung · 2022
Cited alongside, same era.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan · 2024
Later among the works it cites.
Forking paths in neural text generation
Eric Bigelow, Ari Holtzman, Hidenori Tanaka, and Tomer Ullman · 2024
Later among the works it cites.
Deepseek, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Later among the works it cites.
The mystery of the pathological path-star task for language models
Arvid Frydenlund · 2024
Later among the works it cites.
Better & faster large language models via multi-token prediction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen · 2022
Cited alongside, same era.
Diffuseq: Sequence to sequence text generation with diffusion models
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and LingPeng Kong · 2022
Cited alongside, same era.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
Autoregressive modeling with lookahead attention
Li Du, Hongyuan Mei, and Jason Eisner · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Think before you speak: Training language models with pause tokens
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan · 2023
Cited alongside, same era.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Cited alongside, same era.
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve · 2024
Later among the works it cites.
The factorization curse: Which tokens you predict underlie the reversal curse and more
Ouail Kitouni, Niklas S Nolte, Adina Williams, Michael Rabbat, Diane Bouchacourt, and Mark Ibrahim · 2024
Later among the works it cites.
Rho-1: Not all tokens are what you need
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu, Yelong Shen, Ruochen Xu, Chen Lin, Yujiu Yang, Jian Jiao, Nan Duan, et al · 2024
Later among the works it cites.
The clrs-text algorithmic reasoning language benchmark
Larisa Markeeva, Sean McLeish, Borja Ibarz, Wilfried Bounsi, Olga Kozlova, Alex Vitvitskyi, Charles Blundell, Tom Goldstein, Avi Schwarzschild, and Petar Veličković · 2024
Later among the works it cites.
Transformers can navigate mazes with multi-step prediction
Niklas Nolte, Ouail Kitouni, Adina Williams, Mike Rabbat, and Mark Ibrahim · 2024
Later among the works it cites.
σ \sigma -gpts: A new approach to autoregressive models
Arnaud Pannatier, Evann Courdier, and François Fleuret · 2024
Later among the works it cites.
Semformer: Transformer language models with semantic planning
Yongjing Yin, Junran Ding, Kai Song, and Yue Zhang · 2024
Later among the works it cites.
The belief state transformer
Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Dinesh Jayaraman, Alex Lamb, and John Langford · 2025
Closest in time.