Fetching the paper…
Reading the bibliography…
While Reinforcement Learning ( RL) has made great strides towards solving increasingly complicated problems, many algorithms are still brittle to even slight environmental changes.
The algorithm selection problem
J. Rice · 1976
Earlier work this paper cites.
Decision-theoretic planning: Structural assumptions and computational leverage
Craig Boutilier, Thomas L. Dean, and Steve Hanks · 1999
Earlier work this paper cites.
Robust reinforcement learning
J. Morimoto and K. Doya · 2000
Earlier work this paper cites.
Multiagent planning with factored mdps
C. Guestrin, D. Koller, and R. Parr · 2001
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
B. da Silva, E. Basso, A. Bazzan, and P. Engel · 2006
Earlier work this paper cites.
Neuroevolutionary reinforcement learning for generalized helicopter control
R. Koppejan and S. Whiteson · 2009
Earlier work this paper cites.
Abstraction and generalization in reinforcement learning: A summary and framework
M. Ponsen, M. Taylor, and K. Tuyls · 2009
Earlier work this paper cites.
Reinforcement learning-based multi-agent system for network traffic signal control
I. Arel, C. Liu, T. Urbanik, and A. Kohls · 2010
Earlier work this paper cites.
Hydra: Automatically configuring algorithms for portfolio-based selection
L. Xu, H. Hoos, and K. Leyton-Brown · 2010
Earlier work this paper cites.
Protecting against evaluation overfitting in empirical reinforcement learning
S. Whiteson, B. Tanner, M. Taylor, and P. Stone · 2011
Earlier work this paper cites.
Learning parameterized skills
B. da Silva, G. Konidaris, and A. Barto · 2012
Earlier work this paper cites.
Reinforcement learning to adjust parametrized motor primitives to new situations
J. Kober, A. Wilhelm, E. Öztop, and J. Peters · 2012
Earlier work this paper cites.
Contextual markov decision processes
A. Hallak, D. Di Castro, and S. Mannor · 2015
Earlier work this paper cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
F. Doshi-Velez and G. Dimitri Konidaris · 2016
Earlier work this paper cites.
Rl$ˆ2$: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. Bartlett, I. Sutskever, and P. Abbeel · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. Maddison, A. Guez, L. Sifre, G. Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta · 2017
Earlier work this paper cites.
Learning to reinforcement learn
J. Wang, Z. Kurth-Nelson, H. Soyer, J. Leibo, D. Tirumala, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2017
Earlier work this paper cites.
Optimizing chemical reactions with deep reinforcement learning
Z. Zhou, X. Li, and R. Zare · 2017
Earlier work this paper cites.
CAVE: Configuration assessment, visualization and evaluation
A. Biedenkapp, J. Marben, M. Lindauer, and F. Hutter · 2018
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
C. Florensa, D. Held, X. Geng, and P. Abbeel · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Earlier work this paper cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
M. Machado, M. Bellemare, E. Talvitie, J. Veness, M. Hausknecht, and M. Bowling · 2018
Cited alongside, same era.
Markov decision processes with continuous side information
A. Modi, N. Jiang, S. P. Singh, and A. Tewari · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
A. Nichol, J. Achiam, and J. Schulman · 2018
Cited alongside, same era.
Learning goal embeddings via self-play for hierarchical reinforcement learning
S. Sukhbaatar, E. Denton, A. Szlam, and R. Fergus · 2018
Cited alongside, same era.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. Lillicrap, and M. Riedmiller · 2018
A new representation of successor features for transfer across dissimilar environments
M. Abdolshah, H. Le, T. K. George, S. Gupta, S. Rana, and S. Venkatesh · 2021
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. Castro, A. Courville, and M. Bellemare · 2021
Later among the works it cites.
Minimum-delay adaptation in non-stationary reinforcement learning via online high-confidence change-point detection
L. Alegre, A. Bazzan, and B. da Silva · 2021
Later among the works it cites.
P. Castro, T. Kastner, P. Panangaden, and M. Rowland · 2021
Later among the works it cites.
Self-paced context evaluation for contextual reinforcement learning
T. Eimer, A. Biedenkapp, F. Hutter, and M. Lindauer · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hyperparameter importance across datasets
J. van Rijn and F. Hutter · 2018
Cited alongside, same era.
Provably efficient RL with rich observations via latent state decoding
S. Du, A. Krishnamurthy, N. Jiang, A. Agarwal, M. Dudík, and J. Langford · 2019
Cited alongside, same era.
Search on the replay buffer: Bridging planning and reinforcement learning
B. Eysenbach, R. Salakhutdinov, and S. Levine · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning
C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. Bellemare · 2019
Cited alongside, same era.
Obstacle tower: A generalization challenge in vision, control, and planning
A. Juliani, A. Khalifa, V. Berges, J. Harper, E. Teng, H. Henry, A. Crespi, J. Togelius, and D. Lange · 2019
Cited alongside, same era.
A cooperative multi-agent reinforcement learning framework for resource balancing in complex logistics network
X. Li, J. Zhang, J. Bian, Y. Tong, and T. Liu · 2019
Cited alongside, same era.
Active domain randomization
B. Mehta, M. Diaz, F. Golemo, C. Pal, and L. Paull · 2019
Cited alongside, same era.
Sample-efficient automated deep reinforcement learning
J. Franke, G. Köhler, A. Biedenkapp, and F. Hutter · 2021
Later among the works it cites.
Brax - A differentiable physics engine for large scale rigid body simulation
C. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Later among the works it cites.
Why generalization in RL is difficult: Epistemic pomdps and implicit partial observability
Dibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang, Ryan P. Adams, and Sergey Levine · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
R. Kirk, A. Zhang, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
High confidence generalization for reinforcement learning
J. Kostas, Y. Chandak, S. Jordan, G. Theocharous, and P. Thomas · 2021
Later among the works it cites.
URLB: Unsupervised Reinforcement Learning Benchmark
M. Laskin, D. Yarats, H. Liu, K. Lee, A. Zhan, K. Lu, C. Cang, L. Pinto, and P. Abbeel · 2021
Later among the works it cites.
When is generalizable reinforcement learning tractable?
D. Malik, Y. Li, and P. Ravikumar · 2021
Later among the works it cites.
Robots learn increasingly complex tasks with intrinsic motivation and automatic curriculum learning
S. Nguyen, N. Duminy, A. Manoury, D.e Duhaut, and C. Buche · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
M. Samvelyan, R. Kirk, V. Kurin, J. Parker-Holder, M. Jiang, E. Hambro, F. Petroni, H. Küttler, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Toad-gan: a flexible framework for few-shot level generation in token-based games
F. Schubert, M. Awiszus, and B. Rosenhahn · 2021
Later among the works it cites.
Alchemy: A benchmark and analysis toolkit for meta-reinforcement learning agents
J. Wang, M. King, N. Porcel, Z. Kurth-Nelson, T. Zhu, C. Deck, P. Choy, M. Cassin, M. Reynolds, H. Song, G. Buttimore, D. Reichert, N. Rabinowitz, L. Matthey, D. Hassabis, A. Lerchner, and M. Botvinick · 2021
Later among the works it cites.
Reinforcement learning with prototypical representations
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto · 2021
Later among the works it cites.
Learning robust state abstractions for hidden-parameter block mdps
A. Zhang, S. Sodhani, K. Khetarpal, and J. Pineau · 2021
Later among the works it cites.
Robust reinforcement learning on state observations with learned optimal adversary
H. Zhang, H. Chen, D. Boning, and C. Hsieh · 2021
Later among the works it cites.
Automated dynamic algorithm configuration
S. Adriaensen, A. Biedenkapp, G. Shala, N. Awad, T. Eimer, M. Lindauer, and F. Hutter · 2022
Closest in time.
Magnetic control of tokamak plasmas through deep reinforcement learning
J. Degrave, F. Felici, J. Buchli, M. Neunert, B. Tracey, F. Carpanese, T. Ewalds, R. Hafner, A. Abdolmaleki, D. de Las Casas, C. Donner, L. Fritz, C. Galperti, A. Huber, J. Keeling, M. Tsimpoukelli, J. Kay, A. Merle, J. Moret, S. Noury, F. Pesamosca, D. Pfau, O. Sauter, C. Sommariva, S. Coda, B. Duval, A. Fasoli, P. Kohli, K. Kavukcuoglu, D. Hassabis, and M. Riedmiller · 2022
Closest in time.
The impact of task underspecification in evaluating deep reinforcement learning
V. Jayawardana, C. Tang, S. Li, D. Suo, and C. Wu · 2022
Closest in time.
Partially observable markov decision processes and robotics
H. Kurniawati · 2022
Closest in time.
Goal-conditioned reinforcement learning: Problems and solutions
M. Liu, M. Zhu, and W. Zhang · 2022
Closest in time.
Automated reinforcement learning (autorl): A survey and open problems
J. Parker-Holder, R. Rajan, X. Song, A. Biedenkapp, Y. Miao, T. Eimer, B. Zhang, V. Nguyen, R. Calandra, A. Faust, F. Hutter, and M. Lindauer · 2022
Closest in time.
Coax: Plug-n-play reinforcement learning in python with gymnasium and jax
K. Holsheimer, F. Schubert, B. Beilharz, and L. Tsao · 2023
Closest in time.
A survey of zero-shot generalisation in deep reinforcement learning
R. Kirk, A. Zhang, E. Grefenstette, and T. Rocktäschel · 2023
Closest in time.
Polter: Policy trajectory ensemble regularization for unsupervised reinforcement learning
F. Schubert, C. Benjamins, S. Döhler, B. Rosenhahn, and M. Lindauer · 2023
Closest in time.