Fetching the paper…
Reading the bibliography…
Despite many recent works on Mixture of Experts (MoEs) for resource-efficient Transformer language models, existing methods mostly focus on MoEs for feedforward layers.
The meta-pi network: connectionist rapid adaptation for high-performance multi-speaker phoneme recognition
John B. Hampshire II and Alexander H. Waibel · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to recurrent nets
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
The human knowledge compression prize
Marcus Hutter · 2006
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
The goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
ListOps: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R. Bowman · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Fast transformer decoding: One write-head is all you need
Noam Shazeer · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown et al · 2020
Cited alongside, same era.
A mixture of h - 1 heads is better than h heads
Hao Peng, Roy Schwartz, Dianqi Li, and Noah A. Smith · 2020
Unified scaling laws for routed language models
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche, Eliza Rutherford, Tom Hennigan, Matthew Johnson, Katie Millican, Albin Cassirer, Chris Jones, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, and Karen Simonyan · 2022
Later among the works it cites.
On the representation collapse of sparse mixture of experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei · 2022
Later among the works it cites.
Mixture of attention heads: Selecting attention heads per token
Xiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou, Wenge Rong, and Zhang Xiong · 2022
Later among the works it cites.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2022
Later among the works it cites.
In-context learning and induction heads
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman · 2020
Cited alongside, same era.
BASE layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Cited alongside, same era.
GShard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Cited alongside, same era.
RoFormer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Cited alongside, same era.
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Improving transformer with an admixture of attention heads
Tan Nguyen, Tam Nguyen, Hai Do, Khai Nguyen, Vishwanath Saragadam, Minh Pham, Duy Khuong Nguyen, Nhat Ho, and Stanley J. Osher · 2022
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Later among the works it cites.
OpenAI · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott M. Lundberg, Harsha Nori, Hamid Palangi, Marco Túlio Ribeiro, and Yi Zhang · 2023
Closest in time.
llama.cpp
Georgi Gerganov · 2023
Closest in time.
Approximating two-layer feedforward networks for efficient transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2023
Closest in time.
peS2o (Pretraining Efficiently on S2ORC) Dataset
Luca Soldaini and Kyle Lo · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.