Fetching the paper…
Reading the bibliography…
Conditional computation is a popular strategy to make Transformers more efficient.
Reducing Transformer Depth on Demand with Structured Dropout, September 2019
Angela Fan, Edouard Grave, and Armand Joulin · 1909
Earlier work this paper cites.
Depth-Adaptive Transformer, February 2020
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli · 1910
Earlier work this paper cites.
GLU Variants Improve Transformer, February 2020
Noam Shazeer · 2002
Earlier work this paper cites.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding, June 2020
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2006
Earlier work this paper cites.
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Learning Factored Representations in a Deep Mixture of Experts, March 2014
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2014
Earlier work this paper cites.
Conditional Computation in Neural Networks for faster models, January 2016
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2016
Earlier work this paper cites.
BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
Surat Teerapittayanon, Bradley McDanel, and H.T. Kung · 2016
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization, January 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, January 2017
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
SkipNet: Learning Dynamic Routing in Convolutional Networks
Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E. Gonzalez · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization, January 2019
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Root Mean Square Layer Normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Interpreting GPT: the logit lens, August 2020
nostalgebraist · 2020
Earlier work this paper cites.
DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin · 2020
Earlier work this paper cites.
The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Earlier work this paper cites.
CogView: Mastering Text-to-Image Generation via Transformers, November 2021
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang · 2021
Earlier work this paper cites.
Softmax linear units, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, and Ben Mann · 2022
Cited alongside, same era.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, March 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel · 2022
Cited alongside, same era.
Mixture of Attention Heads: Selecting Attention Heads Per Token
Xiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou, Wenge Rong, and Zhang Xiong · 2022
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
Julie Kallini, Shikhar Murty, Christopher D. Manning, Christopher Potts, and Róbert Csordás · 2024
Later among the works it cites.
From Tokens to Words: On the Inner Lexicon of LLMs
Guy Kaplan, Matanel Oren, Yuval Reif, and Roy Schwartz · 2024
Later among the works it cites.
The Remarkable Robustness of LLMs: Stages of Inference?, June 2024
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Later among the works it cites.
Residual Stream Analysis with Multi-Layer SAEs
Tim Lawson, Lucy Farnik, Conor Houghton, and Laurence Aitchison · 2024
Later among the works it cites.
Forgetting Transformer: Softmax Attention with a Forget Gate
Zhixuan Lin, Evgenii Nikishin, Xu He, and Aaron Courville · 2024
Later among the works it cites.
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect, October 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai · 2023
Cited alongside, same era.
Eliciting Latent Predictions from Transformers with the Tuned Lens, November 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Cited alongside, same era.
Finding Neurons in a Haystack: Case Studies with Sparse Probing, June 2023
Wes Gurnee, Neel Nanda, Matthew Pauly, Katherine Harvey, Dmitrii Troitskii, and Dimitris Bertsimas · 2023
Cited alongside, same era.
RoFormer: Enhanced transformer with Rotary Position Embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2023
Cited alongside, same era.
Learning to Skip for Language Modeling, November 2023
Dewen Zeng, Nan Du, Tao Wang, Yuanzhong Xu, Tao Lei, Zhifeng Chen, and Claire Cui · 2023
Cited alongside, same era.
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang · 2024
Cited alongside, same era.
Flex Attention: A Programming Model for Generating Optimized Attention Kernels, December 2024
Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang, and Horace He · 2024
Cited alongside, same era.
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen · 2024
Later among the works it cites.
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
Pit Neitemeier, Björn Deiseroth, Constantin Eichenberg, and Lukas Balles · 2024
Later among the works it cites.
Byte Latent Transformer: Patches Scale Better Than Tokens, December 2024
Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, and Srinivasan Iyer · 2024
Later among the works it cites.
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale, October 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Later among the works it cites.
Skip Transformers: Efficient Inference through Skip-Routing
Matthew Peroni and Dimitris Bertsimas · 2024
Later among the works it cites.
Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, April 2024
David Raposo, Sam Ritter, Blake Richards, Timothy Lillicrap, Peter Conway Humphreys, and Adam Santoro · 2024
Later among the works it cites.
SpaceByte: Towards Deleting Tokenization from Large Language Modeling
Kevin Slagle · 2024
Later among the works it cites.
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
Ziteng Wang, Jun Zhu, and Jianfei Chen · 2024
Later among the works it cites.
Leveraging the true depth of LLMs, February 2025
Ramón Calvo González, Daniele Paliotta, Matteo Pagliardini, Martin Jaggi, and François Fleuret · 2025
Closest in time.
KellerJordan/modded-nanogpt, May 2025
Keller Jordan · 2025
Closest in time.
karpathy/nanoGPT, May 2025
Andrej Karpathy · 2025
Closest in time.
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
Jeonghoon Kim, Byeongchan Lee, Cheonbok Park, Yeontaek Oh, Beomjun Kim, Taehwan Yoo, Seongjin Shin, Dongyoon Han, Jinwoo Shin, and Kang Min Yoo · 2025
Closest in time.
From Bytes to Ideas: Language Modeling with Autoregressive U-Nets, 2025
Mathurin Videau, Badr Youbi Idrissi, Alessandro Leite, Marc Schoenauer, Olivier Teytaud, and David Lopez-Paz · 2025
Closest in time.