Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning (RL) requires a large number of trials to learn a good policy, especially in environments with sparse rewards.
A framework for behavioural claning
Michael Bain and Claude Sommut · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Cooperation under the shadow of the future: experimental evidence from infinitely repeated games
Pedro Dal Bó · 2005
Earlier work this paper cites.
Game theory of mind
Wako Yoshida, Ray J Dolan, and Karl J Friston · 2008
Earlier work this paper cites.
Search-based structured prediction
Hal Daumé, John Langford, and Daniel Marcu · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
A survey of monte carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
On the mixing time and spectral gap for birth and death chains
Guan-Yu Chen and Laurent Saloff-Coste · 2013
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Earlier work this paper cites.
Faulty reward functions in the wild
J. Clark and D. Amodei · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Playing atari games with deep reinforcement learning and human checkpoint replay
Ionel-Alexandru Hosu and Traian Rebedea · 2016
Earlier work this paper cites.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
Max Kleiman-Weiner, Mark K Ho, Joseph L Austerweil, Michael L Littman, and Joshua B Tenenbaum · 2016
Earlier work this paper cites.
Shiv: Reducing supervisor burden in dagger using support vectors for efficient learning from demonstrations in high dimensional state spaces
Michael Laskey, Sam Staszak, Wesley Yu-Shu Hsieh, Jeffrey Mahler, Florian T Pokorny, Anca D Dragan, and Ken Goldberg · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Michel Galley, Jianfeng Gao, and Dan Jurafsky · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Query-efficient imitation learning for end-to-end autonomous driving
Jiakai Zhang and Kyunghyun Cho · 2016
Cited alongside, same era.
Safe and nested subgame solving for imperfect-information games
Noam Brown and Tuomas Sandholm · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Carlos Florensa, David Held, Markus Wulfmeier, and Pieter Abbeel · 2017
Cited alongside, same era.
Playing hard exploration games by watching youtube
Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, and Nando de Freitas · 2018
Closest in time.
Learning from demonstration in the wild
Feryal Behbahani, Kyriacos Shiarlis, Xi Chen, Vitaly Kurin, Sudhanshu Kasewa, Ciprian Stirbu, João Gomes, Supratik Paul, Frans A Oliehoek, João Messias, et al · 2018
Closest in time.
Forward-backward reinforcement learning
Ashley D Edwards, Laura Downs, and James C Davidson · 2018
Closest in time.
Reinforcement learning from imperfect demonstrations
Yang Gao, Ji Lin, Fisher Yu, Sergey Levine, Trevor Darrell, et al · 2018
Closest in time.
Recall traces: Backtracking models for efficient reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, et al · 2017
Cited alongside, same era.
A note on the passage time of finite-state markov chains
Wenming Hong and Ke Zhou · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Julien Perolat, David Silver, Thore Graepel, et al · 2017
Cited alongside, same era.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Cited alongside, same era.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Adam Lerer and Alexander Peysakhovich · 2017
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in no-limit poker
Matej Moravcík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael H. Bowling · 2017
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2017
Cited alongside, same era.
Anirudh Goyal, Philemon Brakel, William Fedus, Timothy Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2018
Closest in time.
Barc: Backward reachability curriculum for robotic reinforcement learning
Boris Ivanovic, James Harrison, Apoorva Sharma, Mo Chen, and Marco Pavone · 2018
Closest in time.
Learning social conventions in markov games
Adam Lerer and Alexander Peysakhovich · 2018
Closest in time.
Solving the rubik’s cube without human knowledge
Stephen McAleer, Forest Agostinelli, Alexander Shmakov, and Pierre Baldi · 2018
Closest in time.
Pommerman
MultiAgentLearning · 2018
Closest in time.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne · 2018
Closest in time.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Alexander Peysakhovich and Adam Lerer · 2018
Closest in time.
Pommerman: A multi-agent playground, 2018
Cinjon Resnick, Wes Eldridge, David Ha, Denny Britz, Jakob Foerster, Julian Togelius, Kyunghyun Cho, and Joan Bruna · 2018
Closest in time.
Learning montezuma’s revenge from a single demonstration
Tim Salimans and Richard Chen · 2018
Closest in time.
Kickstarting deep reinforcement learning
Simon Schmitt, Jonathan J Hudson, Augustin Zidek, Simon Osindero, Carl Doersch, Wojciech M Czarnecki, Joel Z Leibo, Heinrich Kuttler, Andrew Zisserman, Karen Simonyan, et al · 2018
Closest in time.
A hybrid search agent in pommerman
Hongwei Zhou, Yichen Gong, Luvneesh Mugrai, Ahmed Khalifa, Andy Nealen, and Julian Togelius · 2018
Closest in time.
Reinforcement and imitation learning for diverse visuomotor skills
Yuke Zhu, Ziyu Wang, Josh Merel, Andrei A. Rusu, Tom Erez, Serkan Cabi, Saran Tunyasuvunakool, János Kramár, Raia Hadsell, Nando de Freitas, and Nicolas Heess · 2018
Closest in time.