Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) works impressively in some environments and fails catastrophically in others.
Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset, September 2019
Ruohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan, Karl S. Muller, Jake A. Whritner, Luxin Zhang, Mary M. Hayhoe, and Dana H. Ballard · 1903
Earlier work this paper cites.
Provably Efficient Reinforcement Learning with Linear Function Approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 1907
Earlier work this paper cites.
Behaviour Suite for Reinforcement Learning, February 2020
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado Van Hasselt · 1908
Earlier work this paper cites.
MDP Playground: A Design and Debug Testbed for Reinforcement Learning, June 2021
Raghu Rajan, Jessica Lizeth Borja Diaz, Suresh Guttikonda, Fabio Ferreira, André Biedenkapp, Jan Ole von Hartz, and Frank Hutter · 1909
Earlier work this paper cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 1911
Earlier work this paper cites.
Leveraging Procedural Generation to Benchmark Reinforcement Learning, July 2020
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 1912
Earlier work this paper cites.
Expected-outcome: a general model of static evaluation
B. Abramson · 1939
Earlier work this paper cites.
On an iterative technique for Riccati equation computations
D. Kleinman · 1968
Earlier work this paper cites.
On the Convergence of Policy Iteration in Stationary Dynamic Programming
Martin L. Puterman and Shelby L. Brumelle · 1979
Earlier work this paper cites.
Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Monte Carlo Go
Bernd Brügmann · 1993
Earlier work this paper cites.
On-line Policy Improvement using Monte-Carlo Search
Gerald Tesauro and Gregory Galperin · 1996
Earlier work this paper cites.
Rollout Algorithms for Combinatorial Optimization
Dimitri P. Bertsekas, John N. Tsitsiklis, and Cynara Wu · 1997
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
R-MAX - A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Ronen I. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
A Sparse Sampling Algorithm for Near-Optimal Planning in Large Markov Decision Processes
Michael Kearns, Yishay Mansour, and Andrew Y. Ng · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Learning Rates for Q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Monte-Carlo Go Developments
B. Bouzy and B. Helmstetter · 2004
Earlier work this paper cites.
An Adaptive Sampling Algorithm for Solving Markov Decision Processes
Hyeong Soo Chang, Michael C. Fu, Jiaqiao Hu, and Steven I. Marcus · 2005
Earlier work this paper cites.
Monte Carlo Planning in RTS Games
Michael Chung, Michael Buro, and Jonathan Schaeffer · 2005
Cited alongside, same era.
Tree-Based Batch Mode Reinforcement Learning
Damien Ernst, Pierre Geurts, and Louis Wehenkel · 2005
Cited alongside, same era.
Neural Fitted Q Iteration – First Experiences with a Data Efficient Neural Reinforcement Learning Method
Martin Riedmiller · 2005
Cited alongside, same era.
Bandit Based Monte-Carlo Planning
Levente Kocsis and Csaba Szepesvári · 2006
Cited alongside, same era.
Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search
Rémi Coulom · 2007
Cited alongside, same era.
Irina Shevtsova · 2011
Is Q-learning Provably Efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I. Jordan · 2018
Later among the works it cites.
DeepMind Control Suite, January 2018
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller · 2018
Later among the works it cites.
Minimalistic Gridworld Environment for OpenAI Gym, 2018
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Later among the works it cites.
Yao Liu and Emma Brunskill · 2019
Later among the works it cites.
BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning, December 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Playing Atari with Deep Reinforcement Learning, December 2013
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
The Dependence of Effective Planning Horizon on Model Accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Cited alongside, same era.
Frame skip is a powerful parameter for learning to play Atari
Alex Braylan, Mark Hollenbeck, Elliot Meyerson, and Risto Miikkulainen · 2015
Cited alongside, same era.
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio · 2019
Later among the works it cites.
Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design
Michael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre Bayen, Stuart Russell, Andrew Critch, and Sergey Levine · 2020
Later among the works it cites.
Rollout, Policy Iteration, and Distributed Reinforcement Learning
Dimitri Bertsekas · 2020
Later among the works it cites.
Antonin Raffin · 2020
Later among the works it cites.
Sample Efficient Reinforcement Learning In Continuous State Spaces: A Perspective Beyond Linearity
Dhruv Malik, Aldo Pacchiano, Vishwak Srinivasan, and Yuanzhi Li · 2021
Later among the works it cites.
State Entropy Maximization with Random Encoders for Efficient Exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Later among the works it cites.
Bilinear Classes: A Structural Framework for Provable Generalization in RL
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Later among the works it cites.
Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient Algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Later among the works it cites.
Offline RL Without Off-Policy Evaluation, December 2021
David Brandfonbrener, William F. Whitney, Rajesh Ranganath, and Joan Bruna · 2021
Later among the works it cites.
Datasheets for Datasets, December 2021
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Maximilian Ernestus, and Noah Dormann · 2021
Later among the works it cites.
Guarantees for Epsilon-Greedy Reinforcement Learning with Function Approximation
Chris Dann, Yishay Mansour, Mehryar Mohri, Ayush Sekhari, and Karthik Sridharan · 2022
Later among the works it cites.
The Sandbox Environment for Generalizable Agent Research (SEGAR), March 2022
R. Devon Hjelm, Bogdan Mazoure, Florian Golemo, Felipe Frujeri, Mihai Jalobeanu, and Andrey Kolobov · 2022
Later among the works it cites.
Hardness in Markov Decision Processes: Theory and Practice, October 2022
Michelangelo Conserva and Paulo Rauber · 2022
Later among the works it cites.
Lessons from AlphaZero for Optimal, Model Predictive, and Adaptive Control
Dimitri P. Bertsekas · 2022
Later among the works it cites.