Fetching the paper…
Reading the bibliography…
Multi-task learning (MTL) with neural networks leverages commonalities in tasks to improve performance, but often suffers from task interference which reduces the benefits of transfer.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
A computational model of action selection in the basal ganglia. i. a new functional anatomy
Kevin Gurney, Tony J Prescott, and Peter Redgrave · 2001
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G. Barto and Sridhar Mahadevan · 2003
Earlier work this paper cites.
Learning the task allocation game
Sherief Abdallah and Victor Lesser · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Shifting the spotlight of attention: evidence for discrete computations in cognition
Timothy J Buschman and Earl K Miller · 2010
Earlier work this paper cites.
Conditional routing of information to the cortex: A model of the basal ganglia’s role in cognitive coordination
Andrea Stocco, Christian Lebiere, and John R Anderson · 2010
Earlier work this paper cites.
Deep sequential neural network
Ludovic Denoyer and Patrick Gallinari · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deep networks with internal selective attention through feedback connections
Marijn F Stollenga, Jonathan Masci, Faustino Gomez, and Juergen Schmidhuber · 2014
Cited alongside, same era.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Cited alongside, same era.
Expert gate: Lifelong learning with a network of experts
Rahaf Aljundi, Punarjay Chakravarty, and Tinne Tuytelaars · 2016
Pathnet: Evolution channels gradient descent in super neural networks
Chrisantha Fernando, Dylan Banarse, Charles Blundell, Yori Zwols, David Ha, Andrei A. Rusu, Alexander Pritzel, and Daan Wierstra · 2017
Closest in time.
Metacontrol for adaptive imagination-based optimization
Jessica B Hamrick, Andrew J Ballard, Razvan Pascanu, Oriol Vinyals, Nicolas Heess, and Peter W Battaglia · 2017
Closest in time.
Dynamic deep neural networks: Optimizing accuracy-efficiency trade-offs by selective execution
Lanlan Liu and Jia Deng · 2017
Closest in time.
Deciding how to decide: Dynamic routing in artificial neural networks
Mason McGill and Pietro Perona · 2017
Closest in time.
Risto Miikkulainen, Jason Liang, Elliot Meyerson, Aditya Rawal, Dan Fink, Olivier Francon, Bala Raju, Arshak Navruzyan, Nigel Duffy, and Babak Hodjat · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adanet: Adaptive structural learning of artificial neural networks
Corinna Cortes, Xavi Gonzalvo, Vitaly Kuznetsov, Mehryar Mohri, and Scott Yang · 2016
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V Le · 2016
Cited alongside, same era.
Cross-stitch networks for multi-task learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, and Martial Hebert · 2016
Cited alongside, same era.
Correcting forecasts with multifactor neural attention
Matthew Riemer, Aditya Vempaty, Flavio Calmon, Fenno Heath, Richard Hull, and Elham Khabiri · 2016
Cited alongside, same era.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy P. Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2016
Cited alongside, same era.
Designing neural network architectures using reinforcement learning
Bowen Baker, Otkrist Gupta, Nikhil Naik, and Ramesh Raskar · 2017
Cited alongside, same era.
SMASH: one-shot model architecture search through hypernetworks
Andrew Brock, Theodore Lim, James M. Ritchie, and Nick Weston · 2017
Cited alongside, same era.
Closest in time.
Meta networks
Tsendsuren Munkhdalai and Hong Yu · 2017
Closest in time.
Janarthanan Rajendran, P. Prasanna, Balaraman Ravindran, and Mitesh M. Khapra · 2017
Closest in time.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2017
Closest in time.
Sluice networks: Learning what to share between loosely related tasks
Sebastian Ruder, Joachim Bingel, Isabelle Augenstein, and Anders Søgaard · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Closest in time.
Learned optimizers that scale and generalize
Olga Wichrowska, Niru Maheswaranathan, Matthew W Hoffman, Sergio Gomez Colmenarejo, Misha Denil, Nando de Freitas, and Jascha Sohl-Dickstein · 2017
Closest in time.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2017
Closest in time.