Fetching the paper…
Reading the bibliography…
Adaptive curricula in reinforcement learning (RL) have proven effective for producing policies robust to discrepancies between the train and test environment.
Equilibrium points in n-person games
J. F. Nash et al · 1950
Earlier work this paper cites.
The theory of statistical decision
L. J. Savage · 1951
Earlier work this paper cites.
A problem in the sequential design of experiments
R. Bellman · 1956
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Y. Chervonenkis · 1971
Earlier work this paper cites.
Sample selection bias as a specification error
J. J. Heckman · 1979
Earlier work this paper cites.
ALVINN: an autonomous land vehicle in a neural network
D. Pomerleau · 1988
Earlier work this paper cites.
Evolutionary robotics and the radical envelope-of-noise hypothesis
N. Jakobi · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein · 2002
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
M. O. Duff · 2002
Earlier work this paper cites.
Correcting sample selection bias by unlabeled data
J. Huang, A. Gretton, K. Borgwardt, B. Schölkopf, and A. Smola · 2006
Earlier work this paper cites.
Discriminative learning under covariate shift
S. Bickel, M. Brückner, and T. Scheffer · 2009
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
A. Der Kiureghian and O. Ditlevsen · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
I. Osband, D. Russo, and B. V. Roy · 2013
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
Earlier work this paper cites.
OFFER: off-environment reinforcement learning
K. A. Ciosek and S. Whiteson · 2017
Cited alongside, same era.
Consistent on-line off-policy evaluation
A. Hallak and S. Mannor · 2017
Cited alongside, same era.
Teacher-student curriculum learning
T. Matiisen, A. Oliver, T. Cohen, and J. Schulman · 2017
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta · 2017
Cited alongside, same era.
CAD2RL: real single-image flight without a single real image
F. Sadeghi and S. Levine · 2017
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
M. Dennis, N. Jaques, E. Vinitsky, A. Bayen, S. Russell, A. Critch, and S. Levine · 2020
Later among the works it cites.
“Other-play” for zero-shot coordination
H. Hu, A. Lerer, A. Peysakhovich, and J. Foerster · 2020
Later among the works it cites.
The NetHack Learning Environment
H. Küttler, N. Nardelli, A. H. Miller, R. Raileanu, M. Selvatici, E. Grefenstette, and T. Rocktäschel · 2020
Later among the works it cites.
Teacher algorithms for curriculum learning of deep rl in continuously parameterized environments
R. Portelas, C. Colas, K. Hofmann, and P.-Y. Oudeyer · 2020
Later among the works it cites.
Adaptive trade-offs in off-policy learning
M. Rowland, W. Dabney, and R. Munos · 2020
Later among the works it cites.
Neuroevolution of self-interpretable agents
Y. Tang, D. Nguyen, and D. Ha · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel · 2017
Cited alongside, same era.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
C. Florensa, D. Held, X. Geng, and P. Abbeel · 2018
Cited alongside, same era.
Procedural level generation improves generality of deep reinforcement learning
N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi · 2018
Cited alongside, same era.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2018
Cited alongside, same era.
Later among the works it cites.
Enhanced POET: Open-ended reinforcement learning through unbounded invention of learning challenges and their solutions
R. Wang, J. Lehman, A. Rawal, J. Zhi, Y. Li, J. Clune, and K. Stanley · 2020
Later among the works it cites.
RTFM: generalising to new environment dynamics via reading
V. Zhong, T. Rocktäschel, and E. Grefenstette · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
L. M. Zintgraf, K. Shiarlis, M. Igl, S. Schulze, Y. Gal, K. Hofmann, and S. Whiteson · 2020
Later among the works it cites.
Learning with AMIGo: Adversarially motivated intrinsic goals
A. Campero, R. Raileanu, H. Kuttler, J. B. Tenenbaum, T. Rocktäschel, and E. Grefenstette · 2021
Later among the works it cites.
Off-belief learning
H. Hu, A. Lerer, B. Cui, L. Pineda, N. Brown, and J. N. Foerster · 2021
Later among the works it cites.
Replay-guided adversarial environment design
M. Jiang, M. Dennis, J. Parker-Holder, J. Foerster, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Prioritized level replay
M. Jiang, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation, 2021
O. OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. de Oliveira Pinto, A. Paino, H. Noh, L. Weng, Q. Yuan, C. Chu, and W. Zaremba · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
M. Samvelyan, R. Kirk, V. Kurin, J. Parker-Holder, M. Jiang, E. Hambro, F. Petroni, H. Kuttler, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, N. McAleese, N. Bradley-Schmieg, N. Wong, N. Porcel, R. Raileanu, S. Hughes-Fitt, V. Dalibard, and W. M. Czarnecki · 2021
Later among the works it cites.
When do curricula work?
X. Wu, E. Dyer, and B. Neyshabur · 2021
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
A. Zhang, R. T. McAllister, R. Calandra, Y. Gal, and S. Levine · 2021
Later among the works it cites.