Fetching the paper…
Reading the bibliography…
Exploration is a fundamental challenge in reinforcement learning (RL).
Evolutionary principles in self-referential learning. on learning now to learn: The meta-meta-meta…-hook
Schmidhuber, Jurgen · 1987
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Bengio, Samy, Bengio, Yoshua, Cloutier, Jocelyn, and Gecsei, Jan · 1995
Earlier work this paper cites.
Learning to learn
Thrun, Sebastian and Pratt, Lorien · 1998
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, Sepp, Younger, A. Steven, and Conwell, Peter R · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael J. and Singh, Satinder P · 2002
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, Ronen I. and Tennenholtz, Moshe · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Singh, Satinder P., Barto, Andrew G., and Chentanez, Nuttapong · 2004
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, Emanuel · 2006
Earlier work this paper cites.
Learning omnidirectional path following using dimensionality reduction
Kolter, J. Zico and Ng, Andrew Y · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, Alexander L. and Littman, Michael L · 2007
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, Olivier and Li, Lihong · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Lopes, Manuel, Lang, Tobias, Toussaint, Marc, and yves Oudeyer, Pierre · 2012
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Koch, Gregory, Zemel, Richard, and Salakhutdinov, Ruslan · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P., Hunt, Jonathan J., Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, John, Levine, Sergey, Abbeel, Pieter, Jordan, Michael I., and Moritz, Philipp · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, Bradly C., Levine, Sergey, and Abbeel, Pieter · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
Andrychowicz, Marcin, Denil, Misha, Colmenarejo, Sergio Gomez, Hoffman, Matthew W., Pfau, David, Schaul, Tom, and de Freitas, Nando · 2016
One-shot visual imitation learning via meta-learning
Finn, Chelsea, Yu, Tianhe, Zhang, Tianhao, Abbeel, Pieter, and Levine, Sergey · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, Carlos, Duan, Yan, and Abbeel, Pieter · 2017
Later among the works it cites.
Noisy networks for exploration
Fortunato, Meire, Azar, Mohammad Gheshlaghi, Piot, Bilal, Menick, Jacob, Osband, Ian, Graves, Alex, Mnih, Vlad, Munos, Rémi, Hassabis, Demis, Pietquin, Olivier, Blundell, Charles, and Legg, Shane · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, Tuomas, Tang, Haoran, Abbeel, Pieter, and Levine, Sergey · 2017
Later among the works it cites.
Meta-sgd: Learning to learn quickly for few shot learning
Li, Zhenguo, Zhou, Fengwei, Chen, Fei, and Li, Hang · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc G., Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Rémi · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
Houthooft, Rein, Chen, Xi, Chen, Xi, Duan, Yan, Schulman, John, De Turck, Filip, and Abbeel, Pieter · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Osband, Ian, Blundell, Charles, Pritzel, Alexander, and Roy, Benjamin Van · 2016
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Santoro, Adam, Bartunov, Sergey, Botvinick, Matthew, Wierstra, Daan, and Lillicrap, Timothy P · 2016
Cited alongside, same era.
Matching networks for one shot learning
Vinyals, Oriol, Blundell, Charles, Lillicrap, Tim, Kavukcuoglu, Koray, and Wierstra, Daan · 2016
Cited alongside, same era.
Learning to reinforcement learn
Wang, Jane X., Kurth-Nelson, Zeb, Tirumala, Dhruva, Soyer, Hubert, Leibo, Joel Z., Munos, Rémi, Blundell, Charles, Kumaran, Dharshan, and Botvinick, Matthew · 2016
Cited alongside, same era.
Later among the works it cites.
Meta-learning with temporal convolutions
Mishra, Nikhil, Rohaninejad, Mostafa, Chen, Xi, and Abbeel, Pieter · 2017
Later among the works it cites.
Munkhdalai, Tsendsuren and Yu, Hong · 2017
Later among the works it cites.
Parameter space noise for exploration
Plappert, Matthias, Houthooft, Rein, Dhariwal, Prafulla, Sidor, Szymon, Chen, Richard Y., Chen, Xi, Asfour, Tamim, Abbeel, Pieter, and Andrychowicz, Marcin · 2017
Later among the works it cites.
Optimization as a model for few-shot learning
Ravi, Sachin and Larochelle, Hugo · 2017
Later among the works it cites.
Prototypical networks for few-shot learning
Snell, Jake, Swersky, Kevin, and Zemel, Richard S · 2017
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Stadie, Bradly, Yang, Ge, Houthooft, Rein, Chen, Xi, Duan, Yan, Wu, Yuhuai, Abbeel, Pieter, and Sutskever, Ilya · 2017
Later among the works it cites.
Learning to learn: Meta-critic networks for sample efficient learning
Sung, Flood, Zhang, Li, Xiang, Tao, Hospedales, Timothy M., and Yang, Yongxin · 2017
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, Karol, Springenberg, Jost Tobias, Ziyu Wang, Nicolas Heess, and Riedmiller, Martin · 2018
Closest in time.