Fetching the paper…
Reading the bibliography…
What data or environments to use for training to improve downstream performance is a longstanding and very topical question in reinforcement learning.
Rui Wang, Joel Lehman, Jeff Clune, and Kenneth O. Stanley · 1901
Earlier work this paper cites.
Motion planning in dynamic environments using velocity obstacles
Paolo Fiorini and Zvi Shiller · 1998
Earlier work this paper cites.
Robust dynamic programming
Garud N Iyengar · 2005
Earlier work this paper cites.
Robust solutions to markov decision problems with uncertain transition matrices
Laurent El Ghaoui and Arnab Nilim · 2005
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
Coordinating hundreds of cooperative, autonomous vehicles in warehouses
Peter R Wurman, Raffaello D’Andrea, and Mick Mountz · 2008
Earlier work this paper cites.
Distributionally robust markov decision processes
Huan Xu and Shie Mannor · 2010
Earlier work this paper cites.
Reciprocal n-body collision avoidance
Jur Van Den Berg, Stephen J Guy, Ming Lin, and Dinesh Manocha · 2011
Earlier work this paper cites.
Interaction between learning and development
Lev Vygotsky et al · 2011
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Reinforcement learning in robust markov decision processes
Shiau Hong Lim, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Wheeled mobile robotics: from fundamentals towards autonomous systems
Gregor Klancar, Andrej Zdesar, Saso Blazic, and Igor Skrjanc · 2017
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Towards optimally decentralized multi-robot collision avoidance via deep reinforcement learning
Pinxin Long, Tingxiang Fan, Xinyi Liao, Wenxi Liu, Hao Zhang, and Jia Pan · 2018
Earlier work this paper cites.
Teacher algorithms for curriculum learning of deep RL in continuously parameterized environments
Rémy Portelas, Cédric Colas, Katja Hofmann, and Pierre-Yves Oudeyer · 2019
Earlier work this paper cites.
Teacher-student curriculum learning
Tambet Matiisen, Avital Oliver, Taco Cohen, and John Schulman · 2019
Cited alongside, same era.
Multi-agent pathfinding: Definitions, variants, and benchmarks
Roni Stern, Nathan Sturtevant, Ariel Felner, Sven Koenig, Hang Ma, Thayne Walker, Jiaoyang Li, Dor Atzmon, Liron Cohen, TK Kumar, et al · 2019
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Michael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen, Stuart Russell, Andrew Critch, and Sergey Levine · 2020
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy P. Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Cited alongside, same era.
Deep reinforcement learning with robust and smooth policy
Qianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang, and Tuo Zhao · 2020
Cited alongside, same era.
Distributed multi-robot collision avoidance via deep reinforcement learning for navigation in complex scenarios
MAESTRO: open-ended environment design for multi-agent reinforcement learning
Mikayel Samvelyan, Akbir Khan, Michael Dennis, Minqi Jiang, Jack Parker-Holder, Jakob Nicolaus Foerster, Roberta Raileanu, and Tim Rocktäschel · 2023
Later among the works it cites.
Stabilizing unsupervised environment design with a learned adversary
Ishita Mediratta, Minqi Jiang, Jack Parker-Holder, Michael Dennis, Eugene Vinitsky, and Tim Rocktäschel · 2023
Later among the works it cites.
Proximal Curriculum for Reinforcement Learning Agents
Georgios Tzannetos, Bárbara Gomes Ribeiro, Parameswaran Kamalaruban, and Adish Singla · 2023
Later among the works it cites.
XLand-minigrid: Scalable meta-reinforcement learning environments in JAX
Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, Viacheslav Sinii, Artem Agarkov, and Sergey Kolesnikov · 2023
Later among the works it cites.
Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo de Lazcano, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tingxiang Fan, Pinxin Long, Wenxi Liu, and Jia Pan · 2020
Cited alongside, same era.
Deepmnavigate: Deep reinforced multi-robot navigation unifying local & global collision avoidance
Qingyang Tan, Tingxiang Fan, Jia Pan, and Dinesh Manocha · 2020
Cited alongside, same era.
Automatic curriculum learning for deep RL: A short survey
Rémy Portelas, Cédric Colas, Lilian Weng, Katja Hofmann, and Pierre-Yves Oudeyer · 2020
Cited alongside, same era.
Curriculum learning for reinforcement learning domains: A framework and survey
Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E. Taylor, and Peter Stone · 2020
Cited alongside, same era.
Replay-guided adversarial environment design
Minqi Jiang, Michael Dennis, Jack Parker-Holder, Jakob N. Foerster, Edward Grefenstette, and Tim Rocktäschel · 2021
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Cited alongside, same era.
Distributionally robust partially observable markov decision process with moment-based ambiguity
Hideaki Nakao, Ruiwei Jiang, and Siqian Shen · 2021
Cited alongside, same era.
Later among the works it cites.
Drl-vo: Learning to navigate through crowded dynamic scenes using velocity obstacles
Zhanteng Xie and Philip Dames · 2023
Later among the works it cites.
Jaxmarl: Multi-agent rl environments in jax
Alexander Rutherford, Benjamin Ellis, Matteo Gallici, Jonathan Cook, Andrei Lupu, Gardar Ingvarsson, Timon Willi, Akbir Khan, Christian Schroeder de Witt, Alexandra Souly, et al · 2023
Later among the works it cites.
minimax: Efficient baselines for autocurricula in jax
Minqi Jiang, Michael Dennis, Edward Grefenstette, and Tim Rocktäschel · 2023
Later among the works it cites.
Benchmarking reinforcement learning techniques for autonomous navigation
Zifan Xu, Bo Liu, Xuesu Xiao, Anirudh Nair, and Peter Stone · 2023
Later among the works it cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy P. Lillicrap · 2023
Later among the works it cites.
Reward-free curricula for training robust world models
Marc Rigter, Minqi Jiang, and Ingmar Posner · 2023
Later among the works it cites.
On the foundation of distributionally robust reinforcement learning
Shengbo Wang, Nian Si, Jose H. Blanchet, and Zhengyuan Zhou · 2023
Later among the works it cites.
Adversarial model for offline reinforcement learning
Mohak Bhardwaj, Tengyang Xie, Byron Boots, Nan Jiang, and Ching-An Cheng · 2023
Later among the works it cites.
Corruption-robust offline reinforcement learning with general function approximation
Chenlu Ye, Rui Yang, Quanquan Gu, and Tong Zhang · 2023
Later among the works it cites.
Learning perceptual hallucination for multi-robot navigation in narrow hallways
Jin-Soo Park, Xuesu Xiao, Garrett Warnell, Harel Yedidsion, and Peter Stone · 2023
Later among the works it cites.
Refining minimax regret for unsupervised environment design
Michael Beukman, Samuel Coward, Michael Matthews, Mattie Fellows, Minqi Jiang, Michael Dennis, and Jakob Foerster · 2024
Closest in time.
Craftax: A lightning-fast benchmark for open-ended reinforcement learning
Michael Matthews, Michael Beukman, Benjamin Ellis, Mikayel Samvelyan, Matthew Jackson, Samuel Coward, and Jakob Foerster · 2024
Closest in time.
Jaxued: A simple and useable ued library in jax
Samuel Coward, Michael Beukman, and Jakob Foerster · 2024
Closest in time.
Proximal curriculum with task correlations for deep reinforcement learning
Georgios Tzannetos, Parameswaran Kamalaruban, and Adish Singla · 2024
Closest in time.
Dred: Zero-shot transfer in reinforcement learning via data-regularised environment design
Samuel Garcin, James Doran, Shangmin Guo, Christopher G. Lucas, and Stefano V. Albrecht · 2024
Closest in time.