Fetching the paper…
Reading the bibliography…
Major progress on language models (LMs) in recent years has largely resulted from moving away from specialized models designed for specific tasks, to general models based on powerful architectures (e.g.
“Language Models are Few-shot Learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“ImageNet Classification with Deep Convolutional Neural Networks”
Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton · 2012
Earlier work this paper cites.
“Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation”
Yoshua Bengio, Nicholas L\’eonard and Aaron Courville · 2013
Earlier work this paper cites.
“A Clockwork RNN”
Jan Koutnik, Klaus Greff, Faustino Gomez and Juergen Schmidhuber · 2014
Earlier work this paper cites.
“U-Net: Convolutional Networks for Biomedical Image Segmentation”
Olaf Ronneberger, Philipp Fischer and Thomas Brox · 2015
Earlier work this paper cites.
“Neural Machine Translation of Rare Words with Subword Units”
Rico Sennrich, Barry Haddow and Alexandra Birch · 2015
Earlier work this paper cites.
“Root Mean Square Layer Normalization”
Biao Zhang and Rico Sennrich · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”
Denis Paperno, Germ\’an Kruszewski, Angeliki Lazaridou, Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda and Raquel Fern\’andez · 2016
Earlier work this paper cites.
“Synthetic and Natural Noise Both Break Neural Machine Translation”
Yonatan Belinkov and Yonatan Bisk · 2017
Earlier work this paper cites.
“Dilated Recurrent Neural Networks”
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark Hasegawa-Johnson and Thomas Huang · 2017
Earlier work this paper cites.
“Categorical Reparameterization with Gumbel-Softmax”
Eric Jang, Shixiang Gu and Ben Poole · 2017
Earlier work this paper cites.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
“The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”
C Maddison, A Mnih and Y Teh · 2017
Earlier work this paper cites.
“Outrageously Large Neural Networks: The Sparsely-gated Mixture-of-Experts Layer”
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton and Jeff Dean · 2017
Earlier work this paper cites.
“Neural Discrete Representation Learning”
Aaron Van and Oriol Vinyals · 2017
Earlier work this paper cites.
“Attention is All You Need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, ukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Think You Have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick and Oyvind Tafjord · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
“Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering”
Todor Mihaylov, Peter Clark, Tushar Khot and Ashish Sabharwal · 2018
Earlier work this paper cites.
“Language Models are Unsupervised Multitask Learners”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Ilya Sutskever · 2019
Earlier work this paper cites.
“The Bitter Lesson”, 2019
Richard Sutton · 2019
Earlier work this paper cites.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi and Yejin Choi · 2019
Earlier work this paper cites.
“Longformer: The Long-document Transformer”
Iz Beltagy, Matthew Peters and Arman Cohan · 2020
Earlier work this paper cites.
“PIQA: Reasoning About Physical Commonsense in Natural Language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Earlier work this paper cites.
“Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing”
Zihang Dai, Guokun Lai, Yiming Yang and Quoc Le · 2020
Earlier work this paper cites.
“The Pile: An 800GB Dataset of Diverse Text for Language Modeling”
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser and Connor Leahy · 2020
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu and Dario Amodei · 2020
Earlier work this paper cites.
“Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and Francois Fleuret · 2020
Earlier work this paper cites.
“Adv-BERT: BERT is Not Robust on Misspellings! Generating Nature Adversarial Samples on BERT”
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu and Caiming Xiong · 2020
Earlier work this paper cites.
“Feature Learning in Infinite-width Neural Networks”
Greg Yang and Edward Hu · 2020
Earlier work this paper cites.
“Variable-rate Discrete Representation Learning”
Sander Dieleman, Charlie Nash, Jesse Engel and Karen Simonyan · 2021
Earlier work this paper cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold and Sylvain Gelly · 2021
Earlier work this paper cites.
“WinoGrande: An Adversarial Winograd Schema Challenge at Scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 2021
Earlier work this paper cites.
“CharFormer: Fast Character Transformers via Gradient-based Subword Tokenization”
Yi Tay, Vinh Tran, Sebastian Ruder, Jai Gupta, Hyung Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu and Donald Metzler · 2021
Earlier work this paper cites.
“Token Merging: Your VIT but Faster”
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer and Judy Hoffman · 2022
Earlier work this paper cites.
“Canine: Pre-training an Efficient Tokenization-free Encoder for Language Representation”
Jonathan Clark, Dan Garrette, Iulia Turc and John Wieting · 2022
Earlier work this paper cites.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2022
Earlier work this paper cites.
“MANTa: Efficient Gradient-Based Tokenization for End-to-End Robust Language Modeling”
Nathan Godey, Roman Castagn\’e, \’Eric De and Benot Sagot · 2022
Earlier work this paper cites.
“It’s Raw! Audio Generation with State-Space Models”
Karan Goel, Albert Gu, Chris Donahue and Christopher R\’e · 2022
Cited alongside, same era.
“Efficiently Modeling Long Sequences with Structured State Spaces”
Albert Gu, Karan Goel and Christopher R\’e · 2022
Cited alongside, same era.
“Training Compute-Optimal Large Language Models”
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego Casas, Lisa Hendricks, Johannes Welbl and Aidan Clark · 2022
Cited alongside, same era.
“On the SDEs and Scaling Rules for Adaptive Gradient Algorithms”
Sadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi and Sanjeev Arora · 2022
Cited alongside, same era.
“Hierarchical Transformers Are More Efficient Language Models”
Piotr Nawrot, Szymon Tworkowski, Micha Tyrolski, ukasz Kaiser, Yuhuai Wu, Christian Szegedy and Henryk Michalewski · 2022
Cited alongside, same era.
“Repeat After Me: Transformers are Better Than State Space Models at Copying”
Samy Jelassi, David Brandfonbrener, Sham Kakade and Eran Malach · 2024
Later among the works it cites.
“From Digits to Decisions: How Tokenization Impacts Arithmetic in LLMs”, 2024
Garreth Lee, Guilherme Penedo, Leandro von Werra and Thomas Wolf · 2024
Later among the works it cites.
“Zero-Shot Tokenizer Transfer”
Benjamin Minixhofer, Edoardo Ponti and Ivan Vuli\’c · 2024
Later among the works it cites.
“Introducing OpenAI o1-preview”, 2024
OpenAI · 2024
Later among the works it cites.
“Byte Latent Transformer: Patches Scale Better than Tokens”
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston and Luke Zettlemoyer · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Chain-of-thought Prompting Elicits Reasoning in Large Language Models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed. Chi, Quoc. Le and Denny Zhou · 2022
Cited alongside, same era.
“Byt5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models”
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts and Colin Raffel · 2022
Cited alongside, same era.
“Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models”
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith and Yulia Tsvetkov · 2023
Cited alongside, same era.
“Accelerating Large Language Model Decoding with Speculative Sampling”
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre and John Jumper · 2023
Cited alongside, same era.
“Toucan: Token-aware Character Level Language Modeling”
William Fleshman and Benjamin Van · 2023
Cited alongside, same era.
“Modeling Sequences with Structured State Spaces”, 2023
Albert Gu · 2023
Cited alongside, same era.
“Fast Inference from Transformers via Speculative Decoding”
Yaniv Leviathan, Matan Kalman and Yossi Matias · 2023
Cited alongside, same era.
“The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale”
Guilherme Penedo, Hynek Kydl\’cek, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von and Thomas Wolf · 2024
Later among the works it cites.
“Exact Byte-level Probabilities from Tokenized Language Models for FIM-tasks and Model Ensembles”
Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley and Karen Ullrich · 2024
Later among the works it cites.
“An Analysis of Tokenization: Transformers Under Markov Data”
Nived Rajaraman, Jiantao Jiao and Kannan Ramchandran · 2024
Later among the works it cites.
“Mixture-of-Depths: Dynamically Allocating Compute in Transformer-based Language Models”
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Humphreys and Adam Santoro · 2024
Later among the works it cites.
“Caduceus: Bi-directional Equivariant Long-Range DNA Sequence Modeling”
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao, Albert Gu and Volodymyr Kuleshov · 2024
Later among the works it cites.
“Tokenization Is More Than Compression”
Craig Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine, Omri Uzan, Yuval Pinter and Chris Tanner · 2024
Later among the works it cites.
“SpaceByte: Towards Deleting Tokenization from Large Language Modeling”
Kevin Slagle · 2024
Later among the works it cites.
“Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies”
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin and Ngai Wong · 2024
Later among the works it cites.
“From Language Models over Tokens to Language Models over Characters”
Tim Vieira, Ben LeBrun, Mario Giulianelli, Juan Gastaldi, Brian DuSell, John Terilla, Timothy O’Donnell and Ryan Cotterell · 2024
Later among the works it cites.
“An Empirical Study of Mamba-based Language Models”
Roger Waleffe, Wonmin Byeon, Duncan Riach, Brandon Norick, Vijay Korthikanti, Tri Dao, Albert Gu, Ali Hatamizadeh, Sudhakar Singh and Deepak Narayanan · 2024
Later among the works it cites.
“MambaByte: Token-free Selective State Space Model”
Junxiong Wang, Tushaar Gangavarapu, Jing Yan and Alexander Rush · 2024
Later among the works it cites.
“Gated Linear Attention Transformers with Hardware-Efficient Training”
Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda and Yoon Kim · 2024
Later among the works it cites.
“Parallelizing Linear Transformers with the Delta Rule over Sequence Length”
Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen and Yoon Kim · 2024
Later among the works it cites.
“An Image is Worth 32 Tokens for Reconstruction and Generation”
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers and Liang-Chieh Chen · 2024
Later among the works it cites.
“Genome Modeling and Design Across All Domains of Life with Evo 2”
Garyk Brixi et al · 2025
Closest in time.
Eric Egli, Matteo Manica and Jannis Born · 2025
Closest in time.
“Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach”
Jonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer, Siddharth Singh, Brian. Bartoldson, Bhavya Kailkhura, Abhinav Bhatele and Tom Goldstein · 2025
Closest in time.
“On the Tradeoffs of State Space Models and Transformers”, 2025
Albert Gu · 2025
Closest in time.
Han Guo, Songlin Yang, Tarushii Goel, Eric. Xing, Tri Dao and Yoon Kim · 2025
Closest in time.
“Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling”
Hongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng, Ya Wang, Qiyang Min and Xun Zhou · 2025
Closest in time.
“MrT5: Dynamic Token Merging for Efficient Byte-level Language Models”
Julie Kallini, Shikhar Murty, Christopher Manning, Christopher Potts and R\’obert Csord\’as · 2025
Closest in time.
“Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation”
Shanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang, Yuxi Ren, Xin Xia, Yang Zhao, Xuefeng Xiao and Lu Jiang · 2025
Closest in time.
“SuperBPE: Space Travel for Language Models”
Alisa Liu, Jonathan Hayase, Valentin Hofmann, Sewoong Oh, Noah Smith and Yejin Choi · 2025
Closest in time.
“LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws”
Prasanna Mayilvahanan, Thadd\"aus Wiedemer, Sayak Mallick, Matthias Bethge and Wieland Brendel · 2025
Closest in time.
“Critical Batch Size Revisited: A Simple Empirical Approach to Large-Batch Language Model Training”
William Merrill, Shane Arora, Dirk Groeneveld and Hannaneh Hajishirzi · 2025
Closest in time.
“Universal Cross-Tokenizer Distillation via Approximate Likelihood Matching”
Benjamin Minixhofer, Ivan Vuli\’c and Edoardo Ponti · 2025
Closest in time.
“S1: Simple Test-Time Scaling”
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand\‘es and Tatsunori Hashimoto · 2025
Closest in time.
“The Bitter Lesson is coming for Tokenization”, 2025
Luca Peri\’c · 2025
Closest in time.
“Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles”
Buu Phan, Brandon Amos, Itai Gat, Marton Havasi, Matthew Muckley and Karen Ullrich · 2025
Closest in time.
“From Bytes to Ideas: Language Modeling with Autoregressive U-Nets”
Mathurin Videau, Badr Idrissi, Alessandro Leite, Marc Schoenauer, Olivier Teytaud and David Lopez-Paz · 2025
Closest in time.
“Gated Delta Networks: Improving Mamba2 with Delta Rule”
Songlin Yang, Jan Kautz and Ali Hatamizadeh · 2025
Closest in time.
“Sequential-Parallel Duality in Prefix Scannable Models”
Morris Yau, Sharut Gupta, Valerie Engelmayer, Kazuki Irie, Stefanie Jegelka and Jacob Andreas · 2025
Closest in time.
“OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training”, 2025
Yijiong Yu, Ziyun Dai, Zekun Wang, Wei Wang, Ran Chen and Ji Pei · 2025
Closest in time.