Fetching the paper…
Reading the bibliography…
The Mixture-of-Experts (MoE) architecture is showing promising results in improving parameter sharing in multi-task learning (MTL) and in scaling high-capacity neural networks.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Bias/variance analyses of mixtures-of-experts architectures
Robert A Jacobs · 1997
Earlier work this paper cites.
On the identifiability of mixtures-of-experts
Wenxin Jiang and Martin A Tanner · 1999
Earlier work this paper cites.
Texturing & modeling: a procedural approach
David S Ebert, F Kenton Musgrave, Darwyn Peachey, Ken Perlin, and Steven Worley · 2003
Earlier work this paper cites.
OpenGL shading language
Randi J Rost, Bill Licea-Kane, Dan Ginsburg, John Kessenich, Barthold Lichtenbelt, Hugh Malan, and Mike Weiblen · 2009
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Earlier work this paper cites.
The movielens datasets: History and context
F Maxwell Harper and Joseph A Konstan · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Decision forests, convolutional networks and the models in-between
Yani Ioannou, Duncan Robertson, Darko Zikic, Peter Kontschieder, Jamie Shotton, Matthew Brown, and Antonio Criminisi · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
Andre Martins and Ramon Astudillo · 2016
Cited alongside, same era.
Adaptively sparse transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
Later among the works it cites.
Gumbel-matrix routing for flexible multi-task learning
Krzysztof Maziarz, Efi Kokiopoulou, Andrea Gesmundo, Luciano Sbaiz, Gabor Bartok, and Jesse Berent · 2019
Later among the works it cites.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André FT Martins · 2019
Later among the works it cites.
Reparameterizable subset sampling via continuous relaxations
Sang Michael Xie and Stefano Ermon · 2019
Later among the works it cites.
Recommending what video to watch next: a multitask ranking system
Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed Chi · 2019
Later among the works it cites.
The tree ensemble layer: Differentiability meets conditional computation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sebastian Ruder · 2017
Cited alongside, same era.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin Wainwright, and Michael Jordan · 2018
Cited alongside, same era.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi · 2018
Cited alongside, same era.
Diversity and depth in per-example routing models
Prajit Ramachandran and Quoc V Le · 2018
Cited alongside, same era.
Routing networks: Adaptive selection of non-linear functions for multi-task learning
Clemens Rosenbaum, Tim Klinger, and Matthew Riemer · 2018
Cited alongside, same era.
Hussein Hazimeh, Natalia Ponomareva, Petros Mol, Zhenyu Tan, and Rahul Mazumder · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Later among the works it cites.
Gradient estimation with stochastic softmax tricks
Max Paulus, Dami Choi, Daniel Tarlow, Andreas Krause, and Chris J Maddison · 2020
Later among the works it cites.
Adashare: Learning what to share for efficient deep multi-task learning
Ximeng Sun, Rameswar Panda, Rogerio Feris, and Kate Saenko · 2020
Later among the works it cites.
Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations
Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong · 2020
Later among the works it cites.
Small towers make big differences
Yuyan Wang, Zhe Zhao, Bo Dai, Christopher Fifty, Dong Lin, Lichan Hong, and Ed H Chi · 2020
Later among the works it cites.
Differentiable top-k with optimal transport
Yujia Xie, Hanjun Dai, Minshuo Chen, Bo Dai, Tuo Zhao, Hongyuan Zha, Wei Wei, and Tomas Pfister · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.