Fetching the paper…
Reading the bibliography…
Ability to continuously learn and adapt from limited experience in nonstationary environments is an important milestone on the path towards general intelligence.
Evolutionary principles in self-referential learning
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Learning a synaptic learning rule
Yoshua Bengio, Samy Bengio, and Jocelyn Cloutier · 1990
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Jürgen Schmidhuber · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Continual learning in reinforcement environments
Mark B Ring · 1994
Earlier work this paper cites.
CHILD: A first step towards continual learning
Mark B Ring · 1997
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder Singh, Michael Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Michael Bowling · 2005
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
Bruno C Da Silva, Eduardo W Basso, Ana LC Bazzan, and Paulo M Engel · 2006
Earlier work this paper cites.
Awesome: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
Vincent Conitzer and Tuomas Sandholm · 2007
Cited alongside, same era.
Trueskill™: a bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel · 2007
Cited alongside, same era.
On the role of tracking in stationary environments
Richard S Sutton, Anna Koop, and David Silver · 2007
Cited alongside, same era.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor R Lesser · 2010
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2013
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky · 2016
Later among the works it cites.
Ke Li and Jitendra Malik · 2016
Later among the works it cites.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Later among the works it cites.
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Opponent modelling by sequence prediction and lookahead in two-player games
Richard Mealing and Jonathan L Shapiro · 2013
Cited alongside, same era.
Lifelong machine learning systems: Beyond learning algorithms
Daniel L Silver, Qiang Yang, and Lianghao Li · 2013
Cited alongside, same era.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Cited alongside, same era.
Never ending learning
Tom M Mitchell, William W Cohen, Estevam R Hruschka Jr, Partha Pratim Talukdar, Justin Betteridge, Andrew Carlson, Bhavana Dalvi Mishra, Matthew Gardner, Bryan Kisiel, Jayant Krishnamurthy, et al · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Marcin Andrychowicz, Misha Denil, Sergio Gomez, Matthew W Hoffman, David Pfau, Tom Schaul, and Nando de Freitas · 2016
Cited alongside, same era.
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al · 2016
Later among the works it cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Later among the works it cites.
Contextual explanation networks
Maruan Al-Shedivat, Avinava Dubey, and Eric P Xing · 2017
Closest in time.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Closest in time.
Multiagent bidirectionally-coordinated nets for learning to play starcraft combat games
Peng Peng, Quan Yuan, Ying Wen, Yaodong Yang, Zhenkun Tang, Haitao Long, and Jun Wang · 2017
Closest in time.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Closest in time.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard S Zemel · 2017
Closest in time.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Closest in time.
Recasting gradient-based meta-learning as hierarchical bayes
Erin Grant, Chelsea Finn, Sergey Levine, Trevor Darrell, and Thomas Griffiths · 2018
Closest in time.