Fetching the paper…
Reading the bibliography…
Despite recent success of deep network-based Reinforcement Learning (RL), it remains elusive to achieve human-level efficiency in learning novel tasks.
A Bayesian framework for reinforcement learning. In ICML
Malcolm Strens. 2000 · 2000
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop. 2006 · 2006
Earlier work this paper cites.
Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
Emanuel Todorov, Tom Erez, and Yuval Tassa. 2012 · 2012
Earlier work this paper cites.
Spectral networks and locally connected networks on graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
(More) efficient reinforcement learning via posterior sampling. In Advances in Neural Information Processing Systems
Ian Osband, Daniel Russo, and Benjamin Van Roy. 2013 · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. 2015 · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
Order matters: Sequence to sequence for sets
Oriol Vinyals, Samy Bengio, and Manjunath Kudlur. 2015 · 2015
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics. In Advances in neural information processing systems
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, et al · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Rl2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel. 2016 · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN. In Advances in neural information processing systems
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. 2016 · 2016
Cited alongside, same era.
Matching networks for one shot learning. In Advances in neural information processing systems
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, Daan Wierstra, et al · 2016
Cited alongside, same era.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z. Leibo, Rémi Munos, Charles Blundell, Dharshan Kumaran, and Matthew Botvinick. 2016 · 2016
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Neural message passing for quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning-Volume 70
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Later among the works it cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm. 2018 · 2018
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies. In Advances in Neural Information Processing Systems
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018a · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. 2017 · 2017
Cited alongside, same era.
Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems
Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017 · 2017
Cited alongside, same era.
Vain: Attentional multi-agent predictive modeling. In Advances in Neural Information Processing Systems
Yedid Hoshen. 2017 · 2017
Cited alongside, same era.
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017 · 2017
Cited alongside, same era.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction. In International conference on machine learning
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. 2017 · 2017
Cited alongside, same era.
A simple neural network module for relational reasoning. In Advances in neural information processing systems
Adam Santoro, David Raposo, David G Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in neural information processing systems
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Later among the works it cites.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel. 2018 · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Non-local neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018 · 2018
Later among the works it cites.
Learning to explore via meta-policy gradient. In International Conference on Machine Learning
Tianbing Xu, Qiang Liu, Liang Zhao, and Jian Peng. 2018 · 2018
Later among the works it cites.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess. 2019 · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine. 2019 · 2019
Later among the works it cites.
AlphaStar: Mastering the real-time strategy game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, et al · 2019
Later among the works it cites.
LatentGNN: Learning Efficient Non-local Relations for Visual Recognition
Songyang Zhang, Shipeng Yan, and Xuming He. 2019 · 2019
Later among the works it cites.