Fetching the paper…
Reading the bibliography…
In this paper, we define, evaluate, and improve the ``relay-generalization'' performance of reinforcement learning (RL) agents on the out-of-distribution ``controllable'' states.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Eric Brochu, Vlad M. Cora, and Nando de Freitas · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Carla: An open urban driving simulator
Alexey Dosovitskiy, Germán Ros, Felipe Codevilla, Antonio M. López, and Vladlen Koltun · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel · 2017
Earlier work this paper cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Kumar Gupta · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Joshua Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, P. Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Continual reinforcement learning with complex synapses
Christos Kaplanis, Murray Shanahan, and Claudia Clopath · 2018
Earlier work this paper cites.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and P. Abbeel · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Cited alongside, same era.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Xiaodong Song · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and P. Abbeel · 2018
Cited alongside, same era.
Babyai: A platform to study the sample efficiency of grounded language learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio · 2019
Cited alongside, same era.
Exploring restart distributions
Arash Tavakoli, Vitaly Levdik, Riashat Islam, Christopher M. Smith, and Petar Kormushev · 2020
Later among the works it cites.
Robust deep reinforcement learning against adversarial perturbations on state observations
Huan Zhang, Hongge Chen, Chaowei Xiao, Bo Li, Mingyan D. Liu, Duane S. Boning, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Adaptable agent populations via a generative model of policies
Kenneth Derek and Phillip Isola · 2021
Later among the works it cites.
Challenges and countermeasures for adversarial attacks on deep reinforcement learning
Inaam Ilahi, Muhammad Usama, Junaid Qadir, Muhammad Umar Janjua, Ala Al-Fuqaha, Dinh Thai Hoang, and Dusit Niyato · 2021
Later among the works it cites.
A survey of generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktaschel · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karl Cobbe, Oleg Klimov, Christopher Hesse, Taehoon Kim, and John Schulman · 2019
Cited alongside, same era.
Go-explore: a new approach for hard-exploration problems
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O. Stanley, and Jeff Clune · 2019
Cited alongside, same era.
Implementation matters in deep rl: A case study on ppo and trpo
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2019
Cited alongside, same era.
Sim-to-real via sim-to-sim: Data-efficient robotic grasping via randomized-to-canonical adaptation networks
Stephen James, Paul Wohlhart, Mrinal Kalakrishnan, Dmitry Kalashnikov, Alex Irpan, Julian Ibarz, Sergey Levine, Raia Hadsell, and Konstantinos Bousmalis · 2019
Cited alongside, same era.
Teacher algorithms for curriculum learning of deep rl in continuously parameterized environments
Rémy Portelas, Cédric Colas, Katja Hofmann, and Pierre-Yves Oudeyer · 2019
Cited alongside, same era.
Investigating generalisation in continuous deep reinforcement learning
Chenyang Zhao, Olivier Sigaud, Freek Stulp, and Timothy M. Hospedales · 2019
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Michael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen, Stuart J. Russell, Andrew Critch, and Sergey Levine · 2020
Cited alongside, same era.
Adversarial training blocks generalization in neural policies
Ezgi Korkmaz · 2021
Later among the works it cites.
Language conditioned imitation learning over unstructured data
Corey Lynch and Pierre Sermanet · 2021
Later among the works it cites.
The difficulty of passive learning in deep reinforcement learning
Georg Ostrovski, Pablo Samuel Castro, and Will Dabney · 2021
Later among the works it cites.
Hierarchical reinforcement learning: A comprehensive survey
Shubham Pateria, Budhitama Subagdja, Ah-hwee Tan, and Chai Quek · 2021
Later among the works it cites.
Efficient local planning with linear function approximation
Dong Yin, Botao Hao, Yasin Abbasi-Yadkori, Nevena Lazi’c, and Csaba Szepesvari · 2021
Later among the works it cites.
Robust reinforcement learning on state observations with learned optimal adversary
Huan Zhang, Hongge Chen, Duane S. Boning, and Cho-Jui Hsieh · 2021
Later among the works it cites.
Robust stochastic linear contextual bandits under adversarial attacks
Qin Ding, Cho-Jui Hsieh, and James Sharpnack · 2022
Later among the works it cites.
Are alphazero-like agents robust to adversarial perturbations?
Li-Cheng Lan, Huan Zhang, Ti-Rong Wu, Meng-Yu Tsai, I Wu, Cho-Jui Hsieh, et al · 2022
Later among the works it cites.
Sample efficient deep reinforcement learning via local planning
Dong Yin, Sridhar Thiagarajan, Nevena Lazic, Nived Rajaraman, Botao Hao, and Csaba Szepesvari · 2023
Closest in time.