Fetching the paper…
Reading the bibliography…
Lack of reliability is a well-known issue for reinforcement learning (RL) algorithms.
A Hitchhiker’s Guide to Statistical Comparisons of Reinforcement Learning Algorithms
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 1904
Earlier work this paper cites.
Trends and random walks in macroeconomic time series
Charles R Nelson and Charles I Plosser · 1982
Earlier work this paper cites.
Testing for unit roots in autoregressive-moving average models of unknown order
Said E. Said and David A. Dickey · 1984
Earlier work this paper cites.
Bootstrap Methods for Standard Errors, Confidence Intervals, and Other Measures of Statistical Accuracy
B. Efron and R. Tibshirani · 1986
Earlier work this paper cites.
Alternatives to the MedianAbsolute Deviation
Peter J. Rousseeuw and Christophe Croux · 1993
Earlier work this paper cites.
Time Series Analysis
James D. Hamilton · 1994
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Expected Shortfall: A Natural Coherent Alternative to Value at Risk
Carlo Acerbi and Dirk Tasche · 2002
Earlier work this paper cites.
Drawdown measure in portfolio optimization
Alexei Chekhlov, Stanislav Uryasev, and Michael Zabarankin · 2005
Earlier work this paper cites.
Markov Decision Processes with Average-Value-at-Risk criteria
Nicole Bäuerle and Jonathan Ott · 2011
Earlier work this paper cites.
False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant
Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Algorithms for CVaR Optimization in MDPs
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Optimizing the CVaR via Sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Cited alongside, same era.
Reproducibility of Benchmarked Deep Reinforcement Learning Tasks for Continuous Control
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Later among the works it cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Dopamine: A research framework for deep reinforcement learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Later among the works it cites.
How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benchmarking Deep Reinforcement Learning for Continuous Control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
OpenAI Gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
A Distributional Perspective on Reinforcement Learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Noisy Networks for Exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2017
Cited alongside, same era.
Google Vizier: A Service for Black-Box Optimization
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D. Sculley · 2017
Cited alongside, same era.
Deep Reinforcement Learning that Matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Cited alongside, same era.
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Later among the works it cites.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
TF-Agents: A library for reinforcement learning in tensorflow
Sergio Guadarrama, Anoop Korattikara, Pablo Castro Oscar Ramirez, Ethan Holly, Sam Fishman, Ke Wang, Chris Harris Ekaterina Gonina, Vincent Vanhoucke, and Eugene Brevdo · 2018
Later among the works it cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Matteo Hessel and Joseph Modayil · 2018
Later among the works it cites.
Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Deterministic Implementations for Reproducibility in Deep Reinforcement Learning
Prabhat Nagarajan, Garrett Warnell, and Peter Stone · 2018
Later among the works it cites.