Fetching the paper…
Reading the bibliography…
In unsupervised environment design, reinforcement learning agents are trained on environment configurations (levels) generated by an adversary that maximises some objective.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 1901
Earlier work this paper cites.
Equilibrium points in n-person games
Nash, J. F · 1950
Earlier work this paper cites.
Games against nature
Milnor, J · 1951
Earlier work this paper cites.
The theory of statistical decision
Savage, L. J · 1951
Earlier work this paper cites.
Games and Decisions: Introduction and Critical Survey
Luce, R. D. and Raiffa, H · 1957
Earlier work this paper cites.
Regret in decision making under uncertainty
Bell, D. E · 1982
Earlier work this paper cites.
Sequential equilibria
Kreps, D. M. and Wilson, R · 1982
Earlier work this paper cites.
Regret theory: An alternative theory of rational choice under uncertainty
Loomes, G. and Sugden, R · 1982
Earlier work this paper cites.
Perfect bayesian equilibrium and sequential equilibrium
Fudenberg, D. and Tirole, J · 1991
Earlier work this paper cites.
A course in game theory
Osborne, M. J. and Rubinstein, A · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
An introduction to game theory , volume 3
Osborne, M. J. et al · 2004
Earlier work this paper cites.
Robust solutions to markov decision problems with uncertain transition matrices
El Ghaoui, L. and Nilim, A · 2005
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Oudeyer, P.-Y., Kaplan, F., and Hafner, V. V · 2007
Earlier work this paper cites.
Distributionally robust markov decision processes
Xu, H. and Mannor, S · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
Reinforcement learning in robust markov decision processes
Lim, S. H., Xu, H., and Mannor, S · 2013
Earlier work this paper cites.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Cited alongside, same era.
Game theory: Parts i and ii-with 88 solved exercises. an open access textbook
Bonanno, G · 2015
Cited alongside, same era.
An introduction to decision theory
Peterson, M · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2018
Cited alongside, same era.
Distributionally robust partially observable markov decision process with moment-based ambiguity
Nakao, H., Jiang, R., and Shen, S · 2021
Later among the works it cites.
Provably efficient model-free constrained rl with linear function approximation
Ghosh, A., Zhou, X., and Shroff, N · 2022
Later among the works it cites.
Robust markov decision processes: Beyond rectangularity
Goyal, V. and Grand-Clément, J · 2022
Later among the works it cites.
Grounding aleatoric uncertainty for unsupervised environment design
Jiang, M., Dennis, M., Parker-Holder, J., Lupu, A., Küttler, H., Grefenstette, E., Rocktäschel, T., and Foerster, J · 2022
Later among the works it cites.
Discovered policy optimisation
Lu, C., Kuba, J., Letcher, A., Metz, L., Schroeder de Witt, C., and Foerster, J · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J., Grefenstette, E., and Rocktäschel, T · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Intrinsic motivation and automatic curricula via asymmetric self-play
Sukhbaatar, S., Lin, Z., Kostrikov, I., Synnaeve, G., Szlam, A., and Fergus, R · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Teacher-student curriculum learning
Matiisen, T., Oliver, A., Cohen, T., and Schulman, J · 2019
Cited alongside, same era.
Teacher algorithms for curriculum learning of deep RL in continuously parameterized environments
Portelas, R., Colas, C., Hofmann, K., and Oudeyer, P · 2019
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A. M., Russell, S., Critch, A., and Levine, S · 2020
Cited alongside, same era.
”other-play” for zero-shot coordination
Hu, H., Lerer, A., Peysakhovich, A., and Foerster, J. N · 2020
Cited alongside, same era.
Later among the works it cites.
RORL: robust offline reinforcement learning via conservative smoothing
Yang, R., Bai, C., Ma, X., Wang, Z., Zhang, C., and Han, L · 2022
Later among the works it cites.
Adversarial model for offline reinforcement learning
Bhardwaj, M., Xie, T., Boots, B., Jiang, N., and Cheng, C · 2023
Later among the works it cites.
minimax: Efficient baselines for autocurricula in jax
Jiang, M., Dennis, M., Grefenstette, E., and Rocktäschel, T · 2023
Later among the works it cites.
Stabilizing unsupervised environment design with a learned adversary
Mediratta, I., Jiang, M., Parker-Holder, J., Dennis, M., Vinitsky, E., and Rocktäschel, T · 2023
Later among the works it cites.
MAESTRO: open-ended environment design for multi-agent reinforcement learning
Samvelyan, M., Khan, A., Dennis, M., Jiang, M., Parker-Holder, J., Foerster, J. N., Raileanu, R., and Rocktäschel, T · 2023
Later among the works it cites.
A unified approach to reinforcement learning, quantal response equilibria, and two-player zero-sum games
Sokota, S., D’Orazio, R., Kolter, J. Z., Loizou, N., Lanctot, M., Mitliagkas, I., Brown, N., and Kroer, C · 2023
Later among the works it cites.
Human-timescale adaptation in an open-ended task space
Team, A. A., Bauer, J., Baumli, K., Baveja, S., Behbahani, F. M. P., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., Dasagi, V., Gonzalez, L., Gregor, K., Hughes, E., Kashem, S., Loks-Thompson, M., Openshaw, H., Parker-Holder, J., Pathak, S., Nieves, N. P., Rakicevic, N., Rocktäschel, T., Schroecker, Y., Sygnowski, J., Tuyls, K., York, S., Zacherl, A., and Zhang, L · 2023
Later among the works it cites.
On the foundation of distributionally robust reinforcement learning
Wang, S., Si, N., Blanchet, J. H., and Zhou, Z · 2023
Later among the works it cites.
Emergence of maps in the memories of blind navigation agents
Wijmans, E., Savva, M., Essa, I., Lee, S., Morcos, A. S., and Batra, D · 2023
Later among the works it cites.
Corruption-robust offline reinforcement learning with general function approximation
Ye, C., Yang, R., Gu, Q., and Zhang, T · 2023
Later among the works it cites.
Jaxued: A simple and useable ued library in jax
Coward, S., Beukman, M., and Foerster, J · 2024
Closest in time.
Dred: Zero-shot transfer in reinforcement learning via data-regularised environment design
Garcin, S., Doran, J., Guo, S., Lucas, C. G., and Albrecht, S. V · 2024
Closest in time.