Fetching the paper…
Reading the bibliography…
While Reinforcement Learning has made great strides towards solving ever more complicated tasks, many algorithms are still brittle to even slight changes in their environment.
MDP Playground: Controlling dimensions of hardness in reinforcement learning
Rajan, R., Diaz, J. L. B., Guttikonda, S., Ferreira, F., Biedenkapp, A., and Hutter, F. (2019) · 1909
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J. (2019) · 1912
Earlier work this paper cites.
The algorithm selection problem
Rice, J. (1976) · 1976
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K. (2000) · 2000
Earlier work this paper cites.
Reinforcement learning-based multi-agent system for network traffic signal control
Arel, I., Liu, C., Urbanik, T., and Kohls, A. (2010) · 2010
Earlier work this paper cites.
Hydra: Automatically configuring algorithms for portfolio-based selection
Xu, L., Hoos, H., and Leyton-Brown, K. (2010) · 2010
Earlier work this paper cites.
Contextual markov decision processes
Hallak, A., Castro, D. D., and Mannor, S. (2015) · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P., Sutskever, I., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016) · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P. (2016) · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C., Guez, A., Sifre, L., Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S. (2017) · 2017
Earlier work this paper cites.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K. (2017) · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A. (2017) · 2017
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J., Kurth-Nelson, Z., Soyer, H., Leibo, J., Tirumala, D., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M. (2017) · 2017
Earlier work this paper cites.
Optimizing chemical reactions with deep reinforcement learning
Zhou, Z., Li, X., and Zare, R. (2017) · 2017
Earlier work this paper cites.
CAVE: Configuration assessment, visualization and evaluation
Biedenkapp, A., Marben, J., Lindauer, M., and Hutter, F. (2018) · 2018
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S. (2018) · 2018
Earlier work this paper cites.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K. (2018) · 2018
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P. (2018) · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M., Bellemare, M., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2018) · 2018
Cited alongside, same era.
Markov decision processes with continuous side information
Modi, A., Jiang, N., Singh, S. P., and Tewari, A. (2018) · 2018
Cited alongside, same era.
Hyperparameter importance across datasets
van Rijn, J. and Hutter, F. (2018) · 2018
Cited alongside, same era.
Provably efficient RL with rich observations via latent state decoding
Du, S., Krishnamurthy, A., Jiang, N., Agarwal, A., Dudík, M., and Langford, J. (2019) · 2019
Cited alongside, same era.
A cooperative multi-agent reinforcement learning framework for resource balancing in complex logistics network
Learning quadrupedal locomotion over challenging terrain
Lee, J., Hwangbo, J., Wellhausen, L., Koltun, V., and Hutter, M. (2020) · 2020
Later among the works it cites.
Is deep reinforcement learning ready for practical applications in healthcare? A sensitivity analysis of duel-ddqn for hemodynamic management in sepsis patients
Lu, M., Shahn, Z., Sow, D., Doshi-Velez, F., and Lehman, L. H. (2020) · 2020
Later among the works it cites.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J. (2020) · 2020
Later among the works it cites.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvári, C., Singh, S., Roy, B. V., Sutton, R. S., Silver, D., and van Hasselt, H. (2020) · 2020
Later among the works it cites.
Rl baselines3 zoo
Raffin, A. (2020) · 2020
Later among the works it cites.
Trajectory-wise multiple choice learning for dynamics generalization in reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, X., Zhang, J., Bian, J., Tong, Y., and Liu, T. (2019) · 2019
Cited alongside, same era.
Reinforcement learning in financial markets
Meng, T. and Khushi, M. (2019) · 2019
Cited alongside, same era.
Stable baselines3
Raffin, A., Hill, A., Ernestus, M., Gleave, A., Kanervisto, A., and Dormann, N. (2019) · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Rakelly, K., Zhou, A., Finn, C., Levine, S., and Quillen, D. (2019) · 2019
Cited alongside, same era.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Ray, A., Achiam, J., and Amodei, D. (2019) · 2019
Cited alongside, same era.
Learning to Design RNA
Runge, F., Stoll, D., Falkner, S., and Hutter, F. (2019) · 2019
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. (2019) · 2019
Cited alongside, same era.
Seo, Y., Lee, K., Clavera, I., Kurutach, T., Shin, J., and Abbeel, P. (2020) · 2020
Later among the works it cites.
A new representation of successor features for transfer across dissimilar environments
Abdolshah, M., Le, H., George, T. K., Gupta, S., Rana, S., and Venkatesh, S. (2021) · 2021
Closest in time.
Evolving reinforcement learning algorithms
Co-Reyes, J. D., Miao, Y., Peng, D., Real, E., Le, Q. V., Levine, S., Lee, H., and Faust, A. (2021) · 2021
Closest in time.
Self-paced context evaluation for contextual reinforcement learning
Eimer, T., Biedenkapp, A., Hutter, F., and Lindauer, M. (2021) · 2021
Closest in time.
Brax - A differentiable physics engine for large scale rigid body simulation
Freeman, C. D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O. (2021) · 2021
Closest in time.
High confidence generalization for reinforcement learning
Kostas, J., Chandak, Y., Jordan, S. M., Theocharous, G., and Thomas, P. (2021) · 2021
Closest in time.
Robots learn increasingly complex tasks with intrinsic motivation and automatic curriculum learning
Nguyen, S., Duminy, N., Manoury, A., Duhaut, D., and Buche, C. (2021) · 2021
Closest in time.
Teachmyagent: a benchmark for automatic curriculum learning in deep RL
Romac, C., Portelas, R., Hofmann, K., and Oudeyer, P. (2021) · 2021
Closest in time.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Kuttler, H., Grefenstette, E., and Rocktäschel, T. (2021) · 2021
Closest in time.
Toad-gan: a flexible framework for few-shot level generation in token-based games
Schubert, F., Awiszus, M., and Rosenhahn, B. (2021) · 2021
Closest in time.
Alchemy: A structured task distribution for meta-reinforcement learning
Wang, J., King, M., Porcel, N., Kurth-Nelson, Z., Zhu, T., Deck, C., Choy, P., Cassin, M., Reynolds, M., Song, H., Buttimore, G., Reichert, D., Rabinowitz, N., Matthey, L., Hassabis, D., Lerchner, A., and Botvinick, M. (2021) · 2021
Closest in time.
Reinforcement learning with prototypical representations
Yarats, D., Fergus, R., Lazaric, A., and Pinto, L. (2021) · 2021
Closest in time.
Learning robust state abstractions for hidden-parameter block mdps
Zhang, A., Sodhani, S., Khetarpal, K., and Pineau, J. (2021a) · 2021
Closest in time.
Robust reinforcement learning on state observations with learned optimal adversary
Zhang, H., Chen, H., Boning, D., and Hsieh, C. (2021b) · 2021
Closest in time.