Fetching the paper…
Reading the bibliography…
Despite recent progress in offline learning, these methods are still trained and tested on the same environment.
Hybrid reinforcement/supervised learning of dialogue policies from fixed data sets
James Henderson, Oliver Lemon, and Kallirroi Georgila · 2008
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade · 2010
Earlier work this paper cites.
Sample-efficient batch reinforcement learning for dialogue management optimization
Olivier Pietquin, Matthieu Geist, Senthilkumar Chandramohan, and Hervé Frezza-Buet · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The atari grand challenge dataset
Vitaly Kurin, Sebastian Nowozin, Katja Hofmann, Lucas Beyer, and Bastian Leibe · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel Todorov, and Sham M. Kakade · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Philip Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, and Emma Brunskill · 2017
Earlier work this paper cites.
Starcraft ii: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexander Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich Küttler, John Agapiou, Julian Schrittwieser, et al · 2017
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Christopher Hesse, Taehoon Kim, and J. Schulman · 2018
Earlier work this paper cites.
Generalization and regularization in dqn
Jesse Farebrother, Marlos C. Machado, and Michael H. Bowling · 2018
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, D. Meger, and Doina Precup · 2018
Earlier work this paper cites.
Illuminating generalization in deep reinforcement learning through procedural level generation
Niels Justesen, Ruben Rodriguez Torrado, Philip Bontrager, Ahmed Khalifa, Julian Togelius, and Sebastian Risi · 2018
Earlier work this paper cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael H. Bowling · 2018
Earlier work this paper cites.
Gotta learn fast: A new benchmark for generalization in rl
Alex Nichol, V. Pfau, Christopher Hesse, O. Klimov, and John Schulman · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, Rafal Józefowicz, Scott Gray, Catherine Olsson, Jakub Pachocki, Michael Petrov, Henrique Pondé de Oliveira Pinto, Jonathan Raiman, Tim Salimans, Jeremy Schlatter, Jonas Schneider, Szymon Sidor, Ilya Sutskever, Jie Tang, Filip Wolski, and Susan Zhang · 2019
Earlier work this paper cites.
Benchmarking open-endedness in minimal criterion coevolution
Jonathan C. Brant and Kenneth O. Stanley · 2019
Earlier work this paper cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gomez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott E. Reed, Rae Jeong, Konrad Zolna, Y. Aytar, D. Budden, Mel Vecerík, Oleg O. Sushkov, David Barker, Jonathan Scholz, Misha Denil, N. D. Freitas, and Ziyun Wang · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 2019
Earlier work this paper cites.
MineRL: A large-scale dataset of Minecraft demonstrations
William H. Guss, Brandon Houghton, Nicholay Topin, Phillip Wang, Cayden Codel, Manuela Veloso, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Generalization in reinforcement learning with selective noise injection and information bottleneck
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Earlier work this paper cites.
Obstacle Tower: A Generalization Challenge in Vision, Control, and Planning
Arthur Juliani, Ahmed Khalifa, Vincent-Pierre Berges, Jonathan Harper, Ervin Teng, Hunter Henry, Adam Crespi, Julian Togelius, and Danny Lange · 2019
Earlier work this paper cites.
Assessing generalization in deep reinforcement learning
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun, and Dawn Song · 2019
Earlier work this paper cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Earlier work this paper cites.
Interference and generalization in temporal difference learning
Emmanuel Bengio, Joelle Pineau, and Doina Precup · 2020
Earlier work this paper cites.
Instance-based generalization in reinforcement learning
Martin Bertran, Natalia Martinez, Mariano Phielipp, and Guillermo Sapiro · 2020
Earlier work this paper cites.
Reinforcement learning generalization with surprise minimization
Jerry Zikun Chen · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman · 2020
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Measuring visual generalization in continuous control from pixels
Jake Grigsby and Yanjun Qi · 2020
Earlier work this paper cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Earlier work this paper cites.
Human-centric dialog training via offline reinforcement learning
N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and R. Fergus · 2020
Cited alongside, same era.
The NetHack Learning Environment
Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Cited alongside, same era.
Reinforcement learning with augmented data
Misha Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Cited alongside, same era.
Network randomization: A simple technique for generalization in deep reinforcement learning
Learning one representation to optimize all rewards
Ahmed Touati and Yann Ollivier · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Denis Yarats, R. Fergus, A. Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Andrea Zanette, Martin J Wainwright, and Emma Brunskill · 2021
Later among the works it cites.
Avalon: A benchmark for rl generalization using procedurally generated worlds
Joshua Albrecht, Abraham Fetterman, Bryden Fogelman, Ellie Kitanidis, Bartosz Wróblewski, Nicole Seo, Michael Rosenthal, Maksis Knutins, Zack Polizzi, James Simon, et al · 2022
Later among the works it cites.
Procedural generalization by planning with self-supervised world models
Ankesh Anand, Jacob C. Walker, Yazhe Li, Eszter Vértes, Julian Schrittwieser, Sherjil Ozair, Theophane Weber, and Jessica B. Hamrick · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Deep reinforcement and infomax learning
Bogdan Mazoure, Remi Tachet des Combes, Thang Long Doan, Philip Bachman, and R Devon Hjelm · 2020
Cited alongside, same era.
Offline meta-reinforcement learning with advantage weighting
E. Mitchell, Rafael Rafailov, X. B. Peng, S. Levine, and Chelsea Finn · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2020
Cited alongside, same era.
Visual transfer for reinforcement learning via wasserstein domain confusion
Josh Roy and G. Konidaris · 2020
Cited alongside, same era.
When does return-conditioned supervised learning work for offline reinforcement learning?
David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche, and Joan Bruna · 2022
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Later among the works it cites.
A study of off-policy learning in environments with procedural content generation
Andy Ehrenberg, Robert Kirk, Minqi Jiang, Edward Grefenstette, and Tim Rocktäschel · 2022
Later among the works it cites.
Dribo: Robust deep reinforcement learning via multi-view information bottleneck
Jiameng Fan and Wenchao Li · 2022
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar · 2022
Later among the works it cites.
Powderworld: A platform for understanding generalization via rich task distributions
Kevin Frans and Phillip Isola · 2022
Later among the works it cites.
Dungeons and data: A large-scale nethack dataset
Eric Hambro, Roberta Raileanu, Danielle Rothermel, Vegard Mella, Tim Rocktäschel, Heinrich Kuttler, and Naila Murray · 2022
Later among the works it cites.
Uncertainty-driven exploration for generalization in reinforcement learning
Yiding Jiang, J Zico Kolter, and Roberta Raileanu · 2022
Later among the works it cites.
Efficient scheduling of data augmentation for deep reinforcement learning
Byungchan Ko and Jungseul Ok · 2022
Later among the works it cites.
Offline q-learning on diverse multi-task data both scales and generalizes
Aviral Kumar, Rishabh Agarwal, Xinyang Geng, G. Tucker, and S. Levine · 2022
Later among the works it cites.
The challenges of exploration for offline reinforcement learning
Nathan Lambert, Markus Wulfmeier, William Whitney, Arunkumar Byravan, Michael Bloesch, Vibhavari Dasagi, Tim Hertweck, and Martin Riedmiller · 2022
Later among the works it cites.
Multi-game decision transformers
Kuang-Huei Lee, Ofir Nachum, Mengjiao Sherry Yang, Lisa Lee, Daniel Freeman, Sergio Guadarrama, Ian Fischer, Winnie Xu, Eric Jang, Henryk Michalewski, et al · 2022
Later among the works it cites.
Learning dynamics and generalization in deep reinforcement learning
Clare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska, and Yarin Gal · 2022
Later among the works it cites.
Improving zero-shot generalization in offline reinforcement learning using generalized similarity functions
Bogdan Mazoure, Ilya Kostrikov, Ofir Nachum, and Jonathan J Tompson · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, Pierre-Luc Bacon, and Aaron Courville · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, J. Foerster, Edward Grefenstette, and Tim Rocktaschel · 2022
Later among the works it cites.
Offline meta-reinforcement learning with online self-supervision
Vitchyr H. Pong, Ashvin V. Nair, Laura M. Smith, Catherine Huang, and Sergey Levine · 2022
Later among the works it cites.
Neorl: A near real-world benchmark for offline reinforcement learning
Rong-Jun Qin, Xingyuan Zhang, Songyi Gao, Xiong-Hui Chen, Zewen Li, Weinan Zhang, and Yang Yu · 2022
Later among the works it cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Later among the works it cites.
Reinforcement learning in robotic applications: a comprehensive survey
Bharat Singh, Rajesh Kumar, and Vinay Pratap Singh · 2022
Later among the works it cites.
Investigating multi-task pretraining and generalization in reinforcement learning
Adrien Ali Taiga, Rishabh Agarwal, Jesse Farebrother, Aaron Courville, and Marc G Bellemare · 2022
Later among the works it cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Later among the works it cites.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
Denis Yarats, David Brandfonbrener, Hao Liu, Michael Laskin, Pieter Abbeel, Alessandro Lazaric, and Lerrel Pinto · 2022
Later among the works it cites.
Offline meta-reinforcement learning for industrial insertion
Tony Zhao, Jianlan Luo, Oleg O. Sushkov, Rugile Pevceviciute, N. Heess, Jonathan Scholz, S. Schaal, and S. Levine · 2022
Later among the works it cites.
Real world offline reinforcement learning with realistic data source
G. Zhou, Liyiming Ke, S. Srinivasa, Abhi Gupta, A. Rajeswaran, and Vikash Kumar · 2022
Later among the works it cites.
Sequence modeling is a robust contender for offline reinforcement learning
Prajjwal Bhargava, Rohan Chitnis, Alborz Geramifard, Shagun Sodhani, and Amy Zhang · 2023
Closest in time.
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian · 2023
Closest in time.
Extreme q-learning: Maxent rl without entropy
Divyansh Garg, Joey Hejna, M. Geist, and S. Ermon · 2023
Closest in time.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Philippe Hansen-Estruch, Ilya Kostrikov, Michael Janner, Jakub Grudzien Kuba, and Sergey Levine · 2023
Closest in time.
A survey of zero-shot generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2023
Closest in time.
Challenges and opportunities in offline reinforcement learning from visual observations
Cong Lu, Philip J Ball, Tim GJ Rudner, Jack Parker-Holder, Michael A Osborne, and Yee Whye Teh · 2023
Closest in time.
Landmark attention: Random-access infinite context length for transformers
Amirkeivan Mohtashami and Martin Jaggi · 2023
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Rafael Figueiredo Prudencio, Marcos ROA Maximo, and Esther Luna Colombini · 2023
Closest in time.
What is essential for unseen goal generalization of offline goal-conditioned rl?
Rui Yang, Lin Yong, Xiaoteng Ma, Hao Hu, Chongjie Zhang, and Tong Zhang · 2023
Closest in time.
Closing the gap between td learning and supervised learning–a generalisation point of view
Raj Ghugare, Matthieu Geist, Glen Berseth, and Benjamin Eysenbach · 2024
Closest in time.