Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Later among the works it cites.
Modular networks: learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber · 2018
Later among the works it cites.
Learning latent permutations with gumbel-sinkhorn networks
Gonzalo Mena, David Belanger, Scott Linderman, and Jasper Snoek · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Stochastic optimization of sorting networks via continuous relaxations
Aditya Grover, Eric Wang, Aaron Zweig, and Stefano Ermon · 2019
Later among the works it cites.
Approximating the permanent by sampling from adaptive partitions
Jonathan Kuck, Tri Dao, Hamid Rezatofighi, Ashish Sabharwal, and Stefano Ermon · 2019
Later among the works it cites.
Sinkhorn autoencoders
Giorgio Patrini, Rianne van den Berg, Patrick Forre, Marcello Carioni, Samarth Bhargav, Max Welling, Tim Genewein, and Frank Nielsen · 2019
Later among the works it cites.
Diversity and depth in per-example routing models
Prajit Ramachandran and Quoc V Le · 2019
Later among the works it cites.
Routing networks and the challenges of modular and compositional computation
Original
Clemens Rosenbaum, Ignacio Cases, Matthew Riemer, and Tim Klinger · 2019
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Original
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Closest in time.
Base layers: Simplifying training of large, sparse models
Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer · 2021
Closest in time.
Scaling vision with sparse mixture of experts
Original
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Closest in time.
Hash layers for large sparse models
Original
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, and Jason Weston · 2021
Closest in time.
Exploring sparse expert models and beyond
Original
An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, et al · 2021
Closest in time.