Fetching the paper…
Reading the bibliography…
We investigate the training of sparse layers that use different parameters for different inputs based on hashing in large Transformer models.
Multilevel adaptive hashing
Andrei Z Broder and Anna R Karlin · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney · 1995
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
Scaling large learning problems with hard parallel mixtures
Ronan Collobert, Yoshua Bengio, and Samy Bengio · 2003
Earlier work this paper cites.
Introduction to algorithms
Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein · 2009
Earlier work this paper cites.
Feature hashing for large scale multitask learning
Kilian Weinberger, Anirban Dasgupta, John Langford, Alex Smola, and Josh Attenberg · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Learning to rank with (a lot of) word features
Bing Bai, Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, and Kilian Weinberger · 2010
Earlier work this paper cites.
Twenty years of mixture of experts
Seniha Esen Yuksel, Joseph N Wilson, and Paul D Gader · 2012
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2013
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2014
Cited alongside, same era.
Generalizing and hybridizing count-based and neural language models
Graham Neubig and Chris Dyer · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Cited alongside, same era.
Hard mixtures of experts for large scale weakly supervised vision
Poly-encoders: Architectures and pre-training strategies for fast and accurate multi-sentence scoring
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sam Gross, Marc’Aurelio Ranzato, and Arthur Szlam · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Efficient softmax approximation for gpus
Armand Joulin, Moustapha Cissé, David Grangier, Hervé Jégou, et al · 2017
Cited alongside, same era.
Lightweight adaptive mixture of neural and n-gram language models
Anton Bakhtin, Arthur Szlam, Marc’Aurelio Ranzato, and Edouard Grave · 2018
Cited alongside, same era.
Learning semantic textual similarity from conversations
Yinfei Yang, Steve Yuan, Daniel Cer, Sheng-yi Kong, Noah Constant, Petr Pilar, Heming Ge, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil · 2018
Cited alongside, same era.
Training millions of personalized dialogue agents
Pierre-Emmanuel Mazaré, Samuel Humeau, Martin Raison, and Antoine Bordes · 2018
Cited alongside, same era.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher · 2019
Cited alongside, same era.
The dialogue dodecathlon: Open-domain knowledge and image grounded conversational agents, 2019
Kurt Shuster, Da Ju, Stephen Roller, Emily Dinan, Y-Lan Boureau, and Jason Weston · 2019
Cited alongside, same era.
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn · 2020
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Base layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Closest in time.