Fetching the paper…
Reading the bibliography…
In order to improve reproducibility, deep reinforcement learning (RL) has been adopting better scientific practices such as standardized evaluation metrics and reporting.
Differential evolution - A simple and efficient heuristic for global optimization over continuous spaces
Storn, R. and Price, K · 1997
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
An efficient approach for assessing hyperparameter importance
Hutter, F., Hoos, H., and Leyton-Brown, K · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M. A., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N · 2016
Earlier work this paper cites.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Ae: A domain-agnostic platform for adaptive experimentation
Bakshy, E., Dworkin, L., Karrer, B., Kashin, K., Letham, B., Murthy, A., and Singh, S · 2018
Earlier work this paper cites.
Minimalistic gridworld environment for gymnasium, 2018
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Earlier work this paper cites.
Neural networks for predicting algorithm runtime distributions
Eggensperger, K., Lindauer, M., and Hutter, F · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2018
Earlier work this paper cites.
Tune: A research platform for distributed model selection and training
Liaw, R., Liang, E., Nishihara, R., Moritz, P., Gonzalez, J., and Stoica, I · 2018
Earlier work this paper cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H., and Silver, D · 2018
Earlier work this paper cites.
Optuna: A next-generation hyperparameter optimization framework
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S · 2019
Earlier work this paper cites.
Pitfalls and best practices in algorithm configuration
Eggensperger, K., Lindauer, M., and Hutter, F · 2019
Earlier work this paper cites.
A generalized framework for population based training
Li, A., Spyra, O., Perel, S., Dalibard, V., Jaderberg, M., Gu, C., Budden, D., Harley, T., and Gupta, P · 2019
Earlier work this paper cites.
Fast efficient hyperparameter tuning for policy gradient methods
Paul, S., Kurin, V., and Whiteson, S · 2019
Earlier work this paper cites.
Hydra - a framework for elegantly configuring complex applications
Yadan, O · 2019
Cited alongside, same era.
Agent57: Outperforming the atari human benchmark
Badia, A., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z., and Blundell, C · 2020
Cited alongside, same era.
Meta learning via learned loss
Bechtle, S., Molchanov, A., Chebotar, Y., Grefenstette, E., Righetti, L., Sukhatme, G., and Meier, F · 2020
Cited alongside, same era.
Dynamic Algorithm Configuration: Foundation of a New Meta-Algorithmic Framework
Biedenkapp, A., Bozkurt, H. F., Eimer, T., Hutter, F., and Lindauer, M · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Cited alongside, same era.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Obando-Ceron, J. and Castro, P · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Later among the works it cites.
Decoupling value and policy for generalization in reinforcement learning
Raileanu, R. and Fergus, R · 2021
Later among the works it cites.
Turner, R., Eriksson, D., McCourt, M., Kiili, J., Laaksonen, E., Xu, Z., and Guyon, I · 2021
Later among the works it cites.
Towards hyperparameter-free policy selection for offline reinforcement learning
Zhang, S. and Jiang, N · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sample-efficient automated deep reinforcement learning
Franke, J. K., Köhler, G., Biedenkapp, A., and Hutter, F · 2020
Cited alongside, same era.
Revisiting design choices in proximal policy optimization
Hsu, C., Mendler-Dünner, C., and Hardt, M · 2020
Cited alongside, same era.
Best practices for scientific research on neural architecture search
Lindauer, M. and Hutter, F · 2020
Cited alongside, same era.
Provably efficient online hyperparameter optimization with population-based bandits
Parker-Holder, J., Nguyen, V., and Roberts, S · 2020
Cited alongside, same era.
RIDE: rewarding impact-driven exploration for procedurally-generated environments
Raileanu, R. and Rocktäschel, T · 2020
Cited alongside, same era.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H., Hessel, M., Oh, J., Singh, S., and Silver, D · 2020
Cited alongside, same era.
Noveld: A simple yet effective exploration criterion
Zhang, T., Xu, H., Wang, X., Wu, Y., Keutzer, K., Gonzalez, J., and Tian, Y · 2021
Later among the works it cites.
Auto-pytorch: Multi-fidelity metalearning for efficient and robust autodl
Zimmer, L., Lindauer, M., and Hutter, F · 2021
Later among the works it cites.
Automated dynamic algorithm configuration
Adriaensen, S., Biedenkapp, A., Shala, G., Awad, N., Eimer, T., Lindauer, M., and Hutter, F · 2022
Later among the works it cites.
Bootstrapped meta-learning
Flennerhag, S., Schroecker, Y., Zahavy, T., van Hasselt, H., Silver, D., and Singh, S · 2022
Later among the works it cites.
Dungeons and data: A large-scale nethack dataset
Hambro, E., Raileanu, R., Rothermel, D., Mella, V., Rocktäschel, T., Küttler, H., and Murray, N · 2022
Later among the works it cites.
Hyperparameter tuning for deep reinforcement learning applications
Kiran, M. and Ozyildirim, B · 2022
Later among the works it cites.
SMAC3: A versatile bayesian optimization package for hyperparameter optimization
Lindauer, M., Eggensperger, K., Feurer, M., Biedenkapp, A., Deng, D., Benjamins, C., Ruhkopf, T., Sass, R., and Hutter, F · 2022
Later among the works it cites.
Discovered policy optimisation
Lu, C., Kuba, J., Letcher, A., Metz, L., de Witt, C., and Foerster, J · 2022
Later among the works it cites.
Automatic termination for hyperparameter optimization
Makarova, A., Shen, H., Perrone, V., Klein, A., Faddoul, J., Krause, A., Seeger, M., and Archambeau, C · 2022
Later among the works it cites.
Velo: Training versatile learned optimizers by scaling up
Metz, L., Harrison, J., Freeman, C., Merchant, A., Beyer, L., Bradbury, J., Agrawal, N., Poole, B., Mordatch, I., Roberts, A., and Sohl-Dickstein, J · 2022
Later among the works it cites.
Automated reinforcement learning (autorl): A survey and open problems
Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., and Lindauer, M · 2022
Later among the works it cites.
Automl loss landscapes
Pushak, Y. and Hoos, H. H · 2022
Later among the works it cites.
Deepcave: An interactive analysis tool for automated machine learning
Sass, R., Bergman, E., Biedenkapp, A., Hutter, F., and Lindauer, M · 2022
Later among the works it cites.
A survey of methods for automated algorithm configuration
Schede, E., Brandt, J., Tornede, A., Wever, M., Bengs, V., Hüllermeier, E., and Tierney, K · 2022
Later among the works it cites.
Autorl-bench 1.0
Shala, G., Arango, S., Biedenkapp, A., Hutter, F., and Grabocka, J · 2022
Later among the works it cites.
Bayesian generational population-based training
Wan, X., Lu, C., Parker-Holder, J., Ball, P., Nguyen, V., Ru, B., and Osborne, M · 2022
Later among the works it cites.