Fetching the paper…
Reading the bibliography…
The combination of Reinforcement Learning (RL) with deep learning has led to a series of impressive feats, with many believing (deep) RL provides a path towards generally capable agents.
Reward shaping via meta-learning
Zou, H., Ren, T., Yan, D., Su, H., and Zhu, J. (2019) · 1901
Earlier work this paper cites.
Deep reinforcement learning using genetic algorithm for parameter optimization
Sehgal, A., La, H. M., Louis, S. J., and Nguyen, H. (2019) · 1905
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L. (2019) · 1910
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S. (2019) · 1912
Earlier work this paper cites.
Adapting behaviour for learning progress
Schaul, T., Borsa, D., Ding, D., Szepesvari, D., Ostrovski, G., Dabney, W., and Osindero, S. (2019) · 1912
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T. P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution.
Rechenberg, I. (1973) · 1973
Earlier work this paper cites.
On bayesian methods for seeking the extremum
Mockus, J. (1974) · 1974
Earlier work this paper cites.
The algorithm selection problem
Rice, J. R. (1976) · 1976
Earlier work this paper cites.
Goal seeking components for adaptive intelligence: An initial assessment.
Barto, A. G., and Sutton, R. S. (1981) · 1981
Earlier work this paper cites.
Model predictive control: Theory and practice - A survey
Garcia, C. E., Prett, D. M., and Morari, M. (1989) · 1989
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S. (1991) · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J., and Dayan, P. (1992) · 1992
Earlier work this paper cites.
On step-size and bias in temporal-difference learning
Sutton, R., and Singh, S. P. (1994) · 1994
Earlier work this paper cites.
Lamarckian evolution, the baldwin effect and function optimization
Whitley, D., Gordon, V. S., and Mathias, K. (1994) · 1994
Earlier work this paper cites.
Adapting crossover in evolutionary algorithms
Spears, W. M. (1995) · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D., and Tsitsiklis, J. (1996) · 1996
Earlier work this paper cites.
Analytical mean squared error curves in temporal difference learning
Singh, S., and Dayan, P. (1996) · 1996
Earlier work this paper cites.
Adaptive critic designs
Prokhorov, D., and Wunsch, D. (1997) · 1997
Earlier work this paper cites.
An overview of parameter control methods by self-adaptation in evolutionary algorithms
Bäck, T. (1998) · 1998
Earlier work this paper cites.
Restart scheduling for genetic algorithms
Fukunaga, A. S. (1998) · 1998
Earlier work this paper cites.
Efficient global optimization of expensive black-box functions
Jones, D. R., Schonlau, M., and Welch, W. J. (1998) · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S. J. (1999) · 1999
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J., and Singh, S. P. (2000) · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. P. (2000) · 2000
Earlier work this paper cites.
Fully differentiable procedural content generation through generative playing networks
Bontrager, P., and Togelius, J. (2020) · 2002
Earlier work this paper cites.
Evolving neural networks through augmenting topologies.
Stanley, K. O., and Miikkulainen, R. (2002) · 2002
Earlier work this paper cites.
Evolution of meta-parameters in reinforcement learning algorithm
Eriksson, A., Capi, G., and Doya, K. (2003) · 2003
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016) · 2003
Earlier work this paper cites.
Self-adaptive evolutionary algorithms.
Gloger, B. (2004) · 2004
Earlier work this paper cites.
Efficient non-linear control through neuroevolution
Gomez, F., Schmidhuber, J., and Miikkulainen, R. (2006) · 2006
Earlier work this paper cites.
Multi-Objective Machine Learning
Jin, Y. (Ed.). (2006) · 2006
Earlier work this paper cites.
Autonomous shaping: knowledge transfer in reinforcement learning
Konidaris, G. D., and Barto, A. G. (2006) · 2006
Earlier work this paper cites.
Automatic data augmentation for generalization in deep reinforcement learning
Raileanu, R., Goldstein, M., Yarats, D., Kostrikov, I., and Fergus, R. (2020) · 2006
Earlier work this paper cites.
Quantity vs. quality: On hyperparameter optimization for deep reinforcement learning
Hertel, L., Baldi, P., and Gillen, D. L. (2020) · 2007
Earlier work this paper cites.
Automatic shaping and decomposition of reward functions.
Marthi, B. (2007) · 2007
Earlier work this paper cites.
Natural selection fails to optimize mutation rates for long-term adaptation on rugged fitness landscapes
Clune, J., Misevic, D., Ofria, C., Lenski, R. E., Elena, S. F., and Sanjuán, R. (2008) · 2008
Earlier work this paper cites.
Evolving coordinated quadruped gaits with the hyperneat generative encoding
Clune, J., Beckmann, B. E., Ofria, C., and Pennock, R. T. (2009) · 2009
Earlier work this paper cites.
Paramils: An automatic algorithm configuration framework
Hutter, F., Hoos, H. H., Leyton-Brown, K., and Stützle, T. (2009) · 2009
Earlier work this paper cites.
A hypercube-based encoding for evolving large-scale neural networks
Stanley, K. O., D’Ambrosio, D. B., and Gauci, J. (2009) · 2009
Earlier work this paper cites.
Generalized domains for empirical evaluations in reinforcement learning.
Whiteson, S., Tanner, B., Taylor, M. E., and Stone, P. (2009) · 2009
Earlier work this paper cites.
Brochu, E., Cora, V. M., and de Freitas, N. (2010) · 2010
Earlier work this paper cites.
Temporal difference bayesian model averaging: A bayesian perspective on adapting lambda
Downey, C., and Sanner, S. (2010) · 2010
Earlier work this paper cites.
ISAC - instance-specific algorithm configuration
Kadioglu, S., Malitsky, Y., Sellmann, M., and Tierney, K. (2010) · 2010
Earlier work this paper cites.
D2RL: deep dense architectures in reinforcement learning
Sinha, S., Bharadhwaj, H., Srinivas, A., and Garg, A. (2020) · 2010
Earlier work this paper cites.
Multi-task evolutionary shaping without pre-specified representations
Snel, M., and Whiteson, S. (2010) · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Srinivas, N., Krause, A., Kakade, S. M., and Seeger, M. W. (2010) · 2010
Earlier work this paper cites.
Hydra: Automatically configuring algorithms for portfolio-based selection
Xu, L., Hoos, H., and Leyton-Brown, K. (2010) · 2010
Earlier work this paper cites.
Griddly: A platform for AI research in games
Bamford, C., Huang, S., and Lucas, S. M. (2020) · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E., and Bach, F. (2011) · 2011
Earlier work this paper cites.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2012) · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J., and Bengio, Y. (2012) · 2012
Earlier work this paper cites.
A survey of monte carlo tree search methods
Browne, C., Powley, E. J., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Liebana, D. P., Samothrakis, S., and Colton, S. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Y., Erez, T., and Todorov, E. (2012) · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
Evolving gaits for physical robots with the hyperneat generative encoding: The benefits of simulation.
Lee, S., Yosinski, J., Glette, K., Lipson, H., and Clune, J. (2013) · 2013
Earlier work this paper cites.
Genetic algorithms+ data structures= evolution programs
Michalewicz, Z. (2013) · 2013
Earlier work this paper cites.
Gaussian processes for nonlinear signal processing: An overview of recent advances
Pérez-Cruz, F., Vaerenbergh, S. V., Murillo-Fuentes, J. J., Lázaro-Gredilla, M., and Santamaría, I. (2013) · 2013
Earlier work this paper cites.
Confronting the challenge of learning a flexible neural controller for a diversity of morphologies
Risi, S., and Stanley, K. O. (2013) · 2013
Earlier work this paper cites.
Learning navigation behaviors end-to-end with autorl
Chiang, H. L., Faust, A., Fiser, M., and Francis, A. G. (2019) · 2014
Earlier work this paper cites.
Reinforcement learning with multi-fidelity simulators
Cutler, M., Walsh, T. J., and How, J. P. (2014) · 2014
Earlier work this paper cites.
A neuroevolution approach to general atari game playing
Hausknecht, M., Lehman, J., Miikkulainen, R., and Stone, P. (2014) · 2014
Earlier work this paper cites.
An efficient approach for assessing hyperparameter importance
Hutter, F., Hoos, H. H., and Leyton-Brown, K. (2014) · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P., and Welling, M. (2014) · 2014
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Moffaert, K. V., and Nowé, A. (2014) · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems.
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X. (2015) · 2015
Earlier work this paper cites.
Bayesian optimization for materials design
Frazier, P. I., and Wang, J. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Self-paced curriculum learning
Jiang, L., Meng, D., Zhao, Q., Shan, S., and Hauptmann, A. G. (2015) · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A., Veness, J., Bellemare, M., Graves, A., Riedmiller, M., Fidjeland, A., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015) · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N. (2015) · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., and de Freitas, N. (2016) · 2016
Earlier work this paper cites.
Openai gym.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Analysing differences between algorithm configurations through ablation
Fawcett, C., and Hoos, H. H. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization
Jamieson, K. G., and Talwalkar, A. (2016) · 2016
Earlier work this paper cites.
Gaussian process bandit optimisation with multi-fidelity evaluations
Kandasamy, K., Dasarathy, G., Oliva, J. B., Schneider, J., and Poczos, B. (2016) · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Earlier work this paper cites.
The whale optimization algorithm
Mirjalili, S., and Lewis, A. (2016) · 2016
Earlier work this paper cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. (2016) · 2016
Earlier work this paper cites.
A greedy approach to adapting the trace parameter for temporal difference learning
White, M., and White, A. (2016) · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Crow, D., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
The option-critic architecture
Bacon, P., Harb, J., and Precup, D. (2017) · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Cited alongside, same era.
Efficient parameter importance analysis via ablation with surrogates
Biedenkapp, A., Lindauer, M., Eggensperger, K., Fawcett, C., Hoos, H. H., and Hutter, F. (2017) · 2017
Cited alongside, same era.
Learning to learn without gradient descent by gradient descent
Chen, Y., Hoffman, M. W., Colmenarejo, S. G., Denil, M., Lillicrap, T. P., Botvinick, M., and de Freitas, N. (2017) · 2017
Botorch: A framework for efficient monte-carlo bayesian optimization
Balandat, M., Karrer, B., Jiang, D. R., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E. (2020) · 2020
Later among the works it cites.
Ready policy one: World building through active learning
Ball, P., Parker-Holder, J., Pacchiano, A., Choromanski, K., and Roberts, S. (2020) · 2020
Later among the works it cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M., Candido, S., Castro, P., Gong, J., Machado, M., Moitra, S., Ponda, S., and Wang, Z. (2020) · 2020
Later among the works it cites.
Basic enhancement strategies when using bayesian optimization for hyperparameter tuning of deep neural networks
Cho, H., Kim, Y., Lee, E., Choi, D., Lee, Y., and Rhee, W. (2020) · 2020
Later among the works it cites.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., and Levine, S. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Online meta-learning by parallel algorithm competition
Elfwing, S., Uchibe, E., and Doya, K. (2017) · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Google vizier: A service for black-box optimization
Golovin, D., Solnik, B., Moitra, S., Kochanski, G., Karro, J., and Sculley, D. (2017) · 2017
Cited alongside, same era.
Parallel and distributed thompson sampling for large-scale accelerated exploration of chemical space
Hernández-Lobato, J. M., Requeima, J., Pyzer-Knapp, E. O., and Aspuru-Guzik, A. (2017) · 2017
Cited alongside, same era.
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
Islam, R., Henderson, P., Gomrokchi, M., and Precup, D. (2017) · 2017
Cited alongside, same era.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K. (2017) · 2017
Cited alongside, same era.
Theory of parameter control for discrete black-box optimization: Provable performance gains through dynamic parameter choices
Doerr, B., and Doerr, C. (2020) · 2020
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. (2020) · 2020
Later among the works it cites.
Growing action spaces
Farquhar, G., Gustafson, L., Lin, Z., Whiteson, S., Usunier, N., and Synnaeve, G. (2020) · 2020
Later among the works it cites.
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W. (2020) · 2020
Later among the works it cites.
Constrained bayesian optimization for automatic chemical design using variational autoencoders
Griffiths, R.-R., and Hernández-Lobato, J. M. (2020) · 2020
Later among the works it cites.
An introduction to surrogate optimization: Intuition, illustration, case study, and the code.
Guo, S. (2020) · 2020
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M. (2020) · 2020
Later among the works it cites.
Learning to utilize shaping rewards: A new approach of reward shaping
Hu, Y., Wang, W., Jia, H., Wang, Y., Chen, Y., Hao, J., Wu, F., and Fan, C. (2020) · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms
Jordan, S., Chandak, Y., Cohen, D., Zhang, M., and Thomas, P. (2020) · 2020
Later among the works it cites.
Model based reinforcement learning for Atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H. (2020) · 2020
Later among the works it cites.
Improving generalization in meta reinforcement learning using learned objectives
Kirsch, L., van Steenkiste, S., and Schmidhuber, J. (2020) · 2020
Later among the works it cites.
Self-paced deep reinforcement learning
Klink, P., D’Eramo, C., Peters, J., and Pajarinen, J. (2020) · 2020
Later among the works it cites.
The NetHack Learning Environment
Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T. (2020) · 2020
Later among the works it cites.
Learning quadrupedal locomotion over challenging terrain
Lee, J., Hwangbo, J., Wellhausen, L., Koltun, V., and Hutter, M. (2020) · 2020
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2020) · 2020
Later among the works it cites.
Curriculum learning for reinforcement learning domains: A framework and survey
Narvekar, S., Peng, B., Leonetti, M., Sinapov, J., Taylor, M. E., and Stone, P. (2020) · 2020
Later among the works it cites.
Knowing the what but not the where in bayesian optimization
Nguyen, V., and Osborne, M. A. (2020) · 2020
Later among the works it cites.
Bayesian optimization for iterative learning
Nguyen, V., Schulze, S., and Osborne, M. A. (2020) · 2020
Later among the works it cites.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Obando-Ceron, J. S., and Castro, P. S. (2020) · 2020
Later among the works it cites.
Discovering reinforcement learning algorithms
Oh, J., Hessel, M., Czarnecki, W. M., Xu, Z., van Hasselt, H., Singh, S., and Silver, D. (2020) · 2020
Later among the works it cites.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvári, C., Singh, S., Van Roy, B., Sutton, R., Silver, D., and van Hasselt, H. (2020) · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, H. F., Rae, J. W., Pascanu, R., Gülçehre, Ç., Jayakumar, S. M., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., Botvinick, M., Heess, N., and Hadsell, R. (2020) · 2020
Later among the works it cites.
Adaptive trade-offs in off-policy learning
Rowland, M., Dabney, W., and Munos, R. (2020) · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al. (2020) · 2020
Later among the works it cites.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B. (2020) · 2020
Later among the works it cites.
Neuroevolution of self-interpretable agents
Tang, Y., Nguyen, D., and Ha, D. (2020) · 2020
Later among the works it cites.
Safe reinforcement learning via curriculum induction
Turchetta, M., Kolobov, A., Shah, S., Krause, A., and Agarwal, A. (2020) · 2020
Later among the works it cites.
Munchausen reinforcement learning
Vieillard, N., Pietquin, O., and Geist, M. (2020) · 2020
Later among the works it cites.
Enhanced POET: open-ended reinforcement learning through unbounded invention of learning challenges and their solutions
Wang, R., Lehman, J., Rawal, A., Zhi, J., Li, Y., Clune, J., and Stanley, K. O. (2020) · 2020
Later among the works it cites.
A performance-based start state curriculum framework for reinforcement learning
Wöhlke, J., Schmitt, F., and van Hoof, H. (2020) · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Xu, Z., van Hasselt, H. P., Hessel, M., Oh, J., Singh, S., and Silver, D. (2020) · 2020
Later among the works it cites.
A self-tuning actor-critic algorithm
Zahavy, T., Xu, Z., Veeriah, V., Hessel, M., Oh, J., van Hasselt, H., Silver, D., and Singh, S. (2020) · 2020
Later among the works it cites.
Automatic curriculum learning through value disagreement
Zhang, Y., Abbeel, P., and Pinto, L. (2020) · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
Zintgraf, L. M., Shiarlis, K., Igl, M., Schulze, S., Gal, Y., Hofmann, K., and Whiteson, S. (2020) · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M. G. (2021) · 2021
Later among the works it cites.
Visionary: Vision architecture discovery for robot learning
Akinola, I., Angelova, A., Lu, Y., Chebotar, Y., Kalashnikov, D., Varley, J., Ibarz, J., and Ryoo, M. S. (2021) · 2021
Later among the works it cites.
What matters for on-policy deep actor-critic methods? a large-scale study
Andrychowicz, M., Raichuk, A., Stańczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., and Bachem, O. (2021) · 2021
Later among the works it cites.
Optimizing hyperparameters of deep reinforcement learning for autonomous driving based on whale optimization algorithm
Ashraf, N. M., Mostafa, R. R., Sakr, R. H., and Rashad, M. Z. (2021) · 2021
Later among the works it cites.
Dehb: Evolutionary hyberband for scalable, robust and efficient hyperparameter optimization
Awad, N., Mallik, N., and Hutter, F. (2021) · 2021
Later among the works it cites.
Logistic q-learning
Bas-Serrano, J., Curi, S., Krause, A., and Neu, G. (2021) · 2021
Later among the works it cites.
Meta learning via learned loss
Bechtle, S., Molchanov, A., Chebotar, Y., Grefenstette, E., Righetti, L., Sukhatme, G. S., and Meier, F. (2020) · 2021
Later among the works it cites.
Carl: A benchmark for contextual and adaptive reinforcement learning
Benjamins, C., Eimer, T., Schubert, F., Biedenkapp, A., Rosenhahn, B., Hutter, F., and Lindauer, M. (2021) · 2021
Later among the works it cites.
TempoRL: Learning when to act
Biedenkapp, A., Rajan, R., Hutter, F., and Lindauer, M. (2021) · 2021
Later among the works it cites.
Learning with {amig}o: Adversarially motivated intrinsic goals
Campero, A., Raileanu, R., Kuttler, H., Tenenbaum, J. B., Rocktäschel, T., and Grefenstette, E. (2021) · 2021
Later among the works it cites.
Evolving reinforcement learning algorithms
Co-Reyes, J. D., Miao, Y., Peng, D., Le, Q. V., Levine, S., Lee, H., and Faust, A. (2021) · 2021
Later among the works it cites.
Temporally-extended ϵ \epsilon -greedy exploration
Dabney, W., Ostrovski, G., and Barreto, A. (2021) · 2021
Later among the works it cites.
Faster improvement rate population based training.
Dalibard, V., and Jaderberg, M. (2021) · 2021
Later among the works it cites.
Hyperparameters in contextual rl are highly situational
Eimer, T., Benjamins, C., and Lindauer, M. (2021a) · 2021
Later among the works it cites.
Adaptive procedural task generation for hard-exploration problems
Fang, K., Zhu, Y., Savarese, S., and Fei-Fei, L. (2021) · 2021
Later among the works it cites.
Learning synthetic environments for reinforcement learning with evolution strategies.
Ferreira, F., Nierhoff, T., and Hutter, F. (2021) · 2021
Later among the works it cites.
Bootstrapped meta-learning
Flennerhag, S., Schroecker, Y., Zahavy, T., van Hasselt, H., Silver, D., and Singh, S. (2021) · 2021
Later among the works it cites.
Sample-efficient automated deep reinforcement learning
Franke, J. K., Koehler, G., Biedenkapp, A., and Hutter, F. (2021) · 2021
Later among the works it cites.
Brax - a differentiable physics engine for large scale rigid body simulation
Freeman, C. D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O. (2021) · 2021
Later among the works it cites.
Quantifying differences in reward functions
Gleave, A., Dennis, M. D., Legg, S., Russell, S., and Leike, J. (2021) · 2021
Later among the works it cites.
Environment generation for zero-shot compositional reinforcement learning
Gur, I., Jaques, N., Miao, Y., Choi, J., Tiwari, M., Lee, H., and Faust, A. (2021) · 2021
Later among the works it cites.
Transient non-stationarity and generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S. (2021) · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Kirk, R., Zhang, A., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Later among the works it cites.
Introducing symmetries to black box meta reinforcement learning
Kirsch, L., Flennerhag, S., van Hasselt, H., Friesen, A. L., Oh, J., and Chen, Y. (2021) · 2021
Later among the works it cites.
Revisiting design choices in offline model based reinforcement learning
Lu, C., Ball, P. J., Parker-Holder, J., Osborne, M., and Roberts, S. (2021) · 2021
Later among the works it cites.
RL-DARTS: differentiable architecture search for reinforcement learning
Miao, Y., Song, X., Peng, D., Yue, S., Brevdo, E., and Faust, A. (2021) · 2021
Later among the works it cites.
Deep reinforcement learning with dynamic optimism
Moskovitz, T., Parker-Holder, J., Pacchiano, A., and Arbel, M. (2021) · 2021
Later among the works it cites.
Deep reinforcement learning for efficient measurement of quantum devices
Nguyen, V., Orbell, S., Lennon, D. T., Moon, H., Vigneau, F., Camenzind, L. C., Yu, L., Zumbühl, D. M., Briggs, G. A. D., Osborne, M. A., et al. (2021) · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation.
OpenAI, O., Plappert, M., Sampedro, R., Xu, T., Akkaya, I., Kosaraju, V., Welinder, P., D’Sa, R., Petron, A., de Oliveira Pinto, H. P., Paino, A., Noh, H., Weng, L., Yuan, Q., Chu, C., and Zaremba, W. (2021) · 2021
Later among the works it cites.
Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRL
Parker-Holder, J., Nguyen, V., Desai, S., and Roberts, S. (2021) · 2021
Later among the works it cites.
The cross-environment hyperparameter setting benchmark for reinforcement learning.
Patterson, A., Neumann, S., White, A. M., Kumaraswamy, R., and White, M. (2021) · 2021
Later among the works it cites.
Amazon sagemaker automatic model tuning: Scalable gradient-free optimization
Perrone, V., Shen, H., Zolic, A., Shcherbatyi, I., Ahmed, A., Bansal, T., Donini, M., Winkelmolen, F., Jenatton, R., Faddoul, J. B., Pogorzelska, B., Miladinovic, M., Kenthapadi, K., Seeger, M. W., and Archambeau, C. (2021) · 2021
Later among the works it cites.
Teachmyagent: a benchmark for automatic curriculum learning in deep RL
Romac, C., Portelas, R., Hofmann, K., and Oudeyer, P. (2021) · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Kuttler, H., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Later among the works it cites.
Song, X., Choromanski, K., Parker-Holder, J., Tang, Y., Peng, D., Jain, D., Gao, W., Pacchiano, A., Sarlós, T., and Yang, Y. (2021) · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
Team, O. E. L., Stooke, A., Mahajan, A., Barros, C., Deck, C., Bauer, J., Sygnowski, J., Trebacz, M., Jaderberg, M., Mathieu, M., McAleese, N., Bradley-Schmieg, N., Wong, N., Porcel, N., Raileanu, R., Hughes-Fitt, S., Dalibard, V., and Czarnecki, W. M. (2021) · 2021
Later among the works it cites.
Simulation-based optimisation to quantify heterogeneity of specific ventilation and perfusion in the lung by the inspired sinewave test
Tran, M., Nguyen, V., Bruce, R., Crockett, D., Formenti, F., Phan, P., Payne, S., and Farmery, A. (2021) · 2021
Later among the works it cites.
Personalized closed-loop brain stimulation for effective neurointervention across participants
van Bueren, N., Reed, T., Nguyen, V., Sheffield, J., van der Ven, S., Osborne, M., Kroesbergen, E., and Kadosh, R. C. (2021) · 2021
Later among the works it cites.
Discovery of options via meta-learned subgoals
Veeriah, V., Zahavy, T., Hessel, M., Xu, Z., Oh, J., Kemaev, I., van Hasselt, H., Silver, D., and Singh, S. (2021) · 2021
Later among the works it cites.
Think global and act local: Bayesian optimisation over high-dimensional categorical and mixed search spaces
Wan, X., Nguyen, V., Ha, H., Ru, B., Lu, C., and Osborne, M. A. (2021) · 2021
Later among the works it cites.
Alchemy: A structured task distribution for meta-reinforcement learning
Wang, J. X., King, M., Porcel, N., Kurth-Nelson, Z., Zhu, T., Deck, C., Choy, P., Cassin, M., Reynolds, M., Song, H. F., Buttimore, G., Reichert, D. P., Rabinowitz, N. C., Matthey, L., Hassabis, D., Lerchner, A., and Botvinick, M. (2021) · 2021
Later among the works it cites.
On the importance of hyperparameter optimization for model-based reinforcement learning
Zhang, B., Rajan, R., Pineda, L., Lambert, N. O., Biedenkapp, A., Chua, K., Hutter, F., and Calandra, R. (2021) · 2021
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de las Casas, D., Donner, C., Fritz, L., Galperti, C., Huber, A., Keeling, J., Tsimpoukelli, M., Kay, J., Merle, A., Moret, J.-M., Noury, S., Pesamosca, F., Pfau, D., Sauter, O., Sommariva, C., Coda, S., Duval, B., Fasoli, A., Kohli, P., Kavukcuoglu, K., Hassabis, D., and Riedmiller, M. (2022) · 2022
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2020) · 2056
Later among the works it cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D. (2016) · 2094
Later among the works it cites.