Fetching the paper…
Reading the bibliography…
Novel reinforcement learning algorithms, or improvements on existing ones, are commonly justified by evaluating their performance on benchmark environments and are compared to an ever-changing set of standard algorithms.
On the “probable error” of a coefficient of correlation deduced from a small sample
Fisher, R. A · 1921
Earlier work this paper cites.
The use of confidence or fiducial limits illustrated in the case of the binomial
Clopper, C. J. and Pearson, E. S · 1934
Earlier work this paper cites.
How not to lie with statistics: The correct way to summarize benchmark results
Fleming, P. J. and Wallace, J. J · 1986
Earlier work this paper cites.
Comparing the areas under two or more correlated receiver operating characteristic curves: a nonparametric approach
DeLong, E. R., DeLong, D. M., and Clarke-Pearson, D. L · 1988
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Efron, B · 1992
Earlier work this paper cites.
Probability Inequalities for sums of Bounded Random Variables , pp. 409–426
Hoeffding, W · 1994
Earlier work this paper cites.
Testing heuristics: We have it all wrong
Hooker, J. N · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
Sutton, R. S · 1995
Earlier work this paper cites.
Reinforcement learning - an introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. P · 1999
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Chentanez, N., Barto, A., and Singh, S · 2004
Earlier work this paper cites.
Utilizing the natural gradient in temporal difference reinforcement learning with eligibility traces
Morimura, T., Uchibe, E., and Doya, K · 2005
Earlier work this paper cites.
Correct equations for the dynamics of the cart-pole system
Florian, R. V · 2007
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Skill discovery in continuous reinforcement learning domains using skill chaining
Konidaris, G. D. and Barto, A. G · 2009
Earlier work this paper cites.
On confidence intervals for P ( X < Y ) P(X<Y)
Kawasaki, Y. and Miyaoka, E · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990-2010)
Schmidhuber, J · 2010
Cited alongside, same era.
Model-free reinforcement learning with continuous action in practice
Degris, T., Pilarski, P. M., and Sutton, R. S · 2012
Cited alongside, same era.
Adaptive step-sizes for reinforcement learning
Dabney, W. C · 2014
Cited alongside, same era.
Bias in natural actor-critic algorithms
Thomas, P · 2014
Cited alongside, same era.
RLPy: A value-function-based reinforcement learning framework for education and research
Geramifard, A., Dann, C., Klein, R. H., Dabney, W., and How, J. P · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Later among the works it cites.
Re-evaluate: Reproducibility in evaluating reinforcement learning algorithms
Smith, J · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
The mirage of action-dependent baselines in reinforcement learning
Tucker, G., Bhupatiraju, S., Gu, S., Turner, R. E., Ghahramani, Z., and Levine, S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D · 2017
Cited alongside, same era.
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Later among the works it cites.
Solving rubik’s cube with a robot hand
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L · 2019
Later among the works it cites.
Confidence intervals for the mann–whitney test
Perme, M. P. and Manevski, D · 2019
Later among the works it cites.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms
Jordan, S. M., Chandak, Y., Cohen, D., Zhang, M., and Thomas, P. S · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2021
Later among the works it cites.
First return, then explore
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2021
Later among the works it cites.
Protecting against evaluation overfitting in empirical reinforcement learning
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P · 2022
Later among the works it cites.
Hyperparameters in reinforcement learning and how to tune them
Eimer, T., Lindauer, M., and Raileanu, R · 2023
Later among the works it cites.
Empirical design in reinforcement learning
Patterson, A., Neumann, S., White, M., and White, A · 2023
Later among the works it cites.