Fetching the paper…
Reading the bibliography…
Some reinforcement learning (RL) algorithms can stitch pieces of experience to solve a task never seen before during training.
Least squares quantization in pcm
S. Lloyd · 1982
Earlier work this paper cites.
How does the value function of a markov decision process depend on the transition probabilities?
Alfred Müller · 1997
Earlier work this paper cites.
Robust reinforcement learning
Jun Morimoto and Kenji Doya · 2000
Earlier work this paper cites.
On the surprising behavior of distance metric in high-dimensional space
Charu Aggarwal, Alexander Hinneburg, and Daniel Keim · 2002
Earlier work this paper cites.
Function approximation via tile coding: Automating parameter choice
Alexander A. Sherstov and Peter Stone · 2005
Earlier work this paper cites.
Lipschitz continuity of value functions in markovian decision processes
K. Hinderer · 2005
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
On the locality of action domination in sequential decision making
Emmanuel Rachelson and Michail Lagoudakis · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Handbook of Markov decision processes: methods and applications
Eugene A Feinberg and Adam Shwartz · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Understanding Machine Learning - From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
The effectiveness of data augmentation in image classification using deep learning, 2017
Luis Perez and Jason Wang · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
A dissection of overfitting and generalization in continuous reinforcement learning, 2018
Amy Zhang, Nicolas Ballas, and Joelle Pineau · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning, 2019
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Earlier work this paper cites.
Generalization in reinforcement learning with selective noise injection and information bottleneck, 2019
Maximilian Igl, Kamil Ciosek, Yingzhen Li, Sebastian Tschiatschek, Cheng Zhang, Sam Devlin, and Katja Hofmann · 2019
Earlier work this paper cites.
Action robust reinforcement learning and applications in continuous control, 2019
Chen Tessler, Yonathan Efroni, and Shie Mannor · 2019
Earlier work this paper cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M. Khoshgoftaar · 2019
Cited alongside, same era.
Learning to reach goals via iterated supervised learning
Dibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu, Coline Devin, Benjamin Eysenbach, and Sergey Levine · 2019
Cited alongside, same era.
Policy continuation with hindsight inverse dynamics
Hao Sun, Zhizhong Li, Xiaotong Liu, Bolei Zhou, and Dahua Lin · 2019
Cited alongside, same era.
Reward-conditioned policies, 2019
Aviral Kumar, Xue Bin Peng, and Sergey Levine · 2019
Cited alongside, same era.
Training neural networks to encode symbols enables combinatorial generalization, 2019
Ivan Vankov and Jeffrey Bowers · 2019
Cited alongside, same era.
Mastering visual continuous control: Improved data-augmented reinforcement learning, 2021
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Multi-game decision transformers, 2022
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski, and Igor Mordatch · 2022
Later among the works it cites.
When does return-conditioned supervised learning work for offline reinforcement learning?
David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche, and Joan Bruna · 2022
Later among the works it cites.
Information prioritization through empowerment in visual model-based rl, 2022
Homanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, and Sergey Levine · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 2019
Cited alongside, same era.
Deepmdp: Learning continuous latent space models for representation learning, 2019
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Cited alongside, same era.
Reinforcement learning upside down: Don’t predict rewards – just map them to actions, 2020
Juergen Schmidhuber · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning, 2020
Aravind Srinivas, Michael Laskin, and Pieter Abbeel · 2020
Cited alongside, same era.
Sample-efficient reinforcement learning via counterfactual-based data augmentation, 2020
Chaochao Lu, Biwei Huang, Ke Wang, José Miguel Hernández-Lobato, Kun Zhang, and Bernhard Schölkopf · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Rvs: What is essential for offline rl via supervised learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine · 2021
Cited alongside, same era.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning, 2022
Denis Yarats, David Brandfonbrener, Hao Liu, Michael Laskin, Pieter Abbeel, Alessandro Lazaric, and Lerrel Pinto · 2022
Later among the works it cites.
Real world offline reinforcement learning with realistic data source, 2022
Gaoyue Zhou, Liyiming Ke, Siddhartha Srinivasa, Abhinav Gupta, Aravind Rajeswaran, and Vikash Kumar · 2022
Later among the works it cites.
Showing your offline reinforcement learning work: Online evaluation budget matters, 2022
Vladislav Kurenkov and Sergey Kolesnikov · 2022
Later among the works it cites.
Bats: Best action trajectory stitching, 2022
Ian Char, Viraj Mehta, Adam Villaflor, John M. Dolan, and Jeff Schneider · 2022
Later among the works it cites.
Imitating past successes can be very suboptimal
Benjamin Eysenbach, Soumith Udatha, Sergey Levine, and Ruslan Salakhutdinov · 2022
Later among the works it cites.
Bisimulation makes analogies in goal-conditioned reinforcement learning, 2022
Philippe Hansen-Estruch, Amy Zhang, Ashvin Nair, Patrick Yin, and Sergey Levine · 2022
Later among the works it cites.
Distinguishing rule- and exemplar-based generalization in learning systems, 2022
Ishita Dasgupta, Erin Grant, and Thomas L. Griffiths · 2022
Later among the works it cites.
Compositional generalization from first principles, 2023
Thaddäus Wiedemer, Prasanna Mayilvahanan, Matthias Bethge, and Wieland Brendel · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought, 2023
Abulhair Saparov and He He · 2023
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task, 2023
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2023
Later among the works it cites.
The benefits of model-based generalization in reinforcement learning, 2023
Kenny Young, Aditya Ramesh, Louis Kirsch, and Jürgen Schmidhuber · 2023
Later among the works it cites.
Pearl: A production-ready reinforcement learning agent, 2023
Zheqing Zhu, Rodrigo de Salvo Braz, Jalaj Bhandari, Daniel Jiang, Yi Wan, Yonathan Efroni, Liyuan Wang, Ruiyang Xu, Hongbo Guo, Alex Nikulkov, Dmytro Korenkevych, Urun Dogan, Frank Cheng, Zheng Wu, and Wanqiao Xu · 2023
Later among the works it cites.
Q-learning decision transformer: Leveraging dynamic programming for conditional sequence modelling in offline RL, 2023
Taku Yamagata, Ahmed Khalil, and Raul Santos-Rodriguez · 2023
Later among the works it cites.
Return augmentation gives supervised RL temporal compositionality, 2023
Keiran Paster, Silviu Pitis, Sheila A. McIlraith, and Jimmy Ba · 2023
Later among the works it cites.
A survey of zero-shot generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2023
Later among the works it cites.
Maxime Chevalier-Boisvert, Bolun Dai, Mark Towers, Rodrigo de Lazcano, Lucas Willems, Salem Lahlou, Suman Pal, Pablo Samuel Castro, and Jordan Terry · 2023
Later among the works it cites.
Gymnasium, March 2023
Mark Towers, Jordan K. Terry, Ariel Kwiatkowski, John U. Balis, Gianluca de Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Tan Jin Shen, and Omar G. Younis · 2023
Later among the works it cites.