Counterfactual data augmentation using locally factored dynamics
Silviu Pitis, Elliot Creager, and Animesh Garg · 2020
Later among the works it cites.
Adaptive trade-offs in off-policy learning
Mark Rowland, Will Dabney, and Rémi Munos · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Reinforcement learning for building controls: The opportunities and challenges
Zhe Wang and Tianzhen Hong · 2020
Later among the works it cites.
Meta-gradient reinforcement learning with an objective discovered online
Zhongwen Xu, Hado P van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Self-tuning deep reinforcement learning
Original
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
Learning retrospective knowledge with reverse reinforcement learning
Shangtong Zhang, Vivek Veeriah, and Shimon Whiteson · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
A survey of exploration methods in reinforcement learning
Original
Susan Amin, Maziar Gomrokchi, Harsh Satija, Herke van Hoof, and Doina Precup · 2021
Later among the works it cites.
An information-theoretic perspective on credit assignment in reinforcement learning
Original
Dilip Arumugam, Peter Henderson, and Pierre-Luc Bacon · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
Parameter-based value functions
Francesco Faccio, Louis Kirsch, and Jürgen Schmidhuber · 2021
Later among the works it cites.
Self-imitation advantage learning
Johan Ferret, Olivier Pietquin, and Matthieu Geist · 2021
Later among the works it cites.
There is no turning back: A self-supervised approach for reversibility-aware reinforcement learning
Nathan Grinsztajn, Johan Ferret, Olivier Pietquin, Matthieu Geist, et al · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Later among the works it cites.
Flexible option learning
Martin Klissarov and Doina Precup · 2021
Later among the works it cites.
Counterfactual credit assignment in model-free reinforcement learning
Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S Stepleton, Nicolas Heess, Arthur Guez, et al · 2021
Later among the works it cites.
Deep reinforcement learning for cyber security
Thanh Thi Nguyen and Vijay Janapa Reddi · 2021
Later among the works it cites.
Posterior value functions: Hindsight baselines for policy gradient methods
Chris Nota, Philip Thomas, and Bruno C. Da Silva · 2021
Later among the works it cites.
Hierarchical reinforcement learning: A comprehensive survey
Shubham Pateria, Budhitama Subagdja, Ah-hwee Tan, and Chai Quek · 2021
Later among the works it cites.
Fitness beats truth in the evolution of perception
Chetan Prakash, Kyle D Stephens, Donald D Hoffman, Manish Singh, and Chris Fields · 2021
Later among the works it cites.
Synthetic returns for long-term credit assignment
David Raposo, Sam Ritter, Adam Santoro, Greg Wayne, Theophane Weber, Matt Botvinick, Hado van Hasselt, and Francis Song · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
Mikayel Samvelyan, Robert Kirk, Vitaly Kurin, Jack Parker-Holder, Minqi Jiang, Eric Hambro, Fabio Petroni, Heinrich Kuttler, Edward Grefenstette, and Tim Rocktäschel · 2021
Later among the works it cites.
Hindsight expectation maximization for goal-conditioned reinforcement learning
Yunhao Tang and Alp Kucukelbir · 2021
Later among the works it cites.
Expected eligibility traces
Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel, David Silver, André Barreto, and Diana Borsa · 2021
Later among the works it cites.
Offline reinforcement learning with reverse model-based imagination
Jianhao Wang, Wenzhe Li, Haozhe Jiang, Guangxiang Zhu, Siyuan Li, and Chongjie Zhang · 2021
Later among the works it cites.
Mastering atari games with limited data
Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, and Yang Gao · 2021
Later among the works it cites.
All you need is supervised learning: From imitation learning to meta-rl with upside down rl
Original
Kai Arulkumaran, Dylan R Ashley, Jürgen Schmidhuber, and Rupesh K Srivastava · 2022
Later among the works it cites.
Learning relative return policies with upside-down reinforcement learning
Original
Dylan R Ashley, Kai Arulkumaran, Jürgen Schmidhuber, and Rupesh Kumar Srivastava · 2022
Later among the works it cites.
On Pearl’s Hierarchy and the Foundations of Causal Inference , pp. 507–556
Elias Bareinboim, Juan D. Correa, Duligur Ibeling, and Thomas Icard · 2022
Later among the works it cites.
Selective credit assignment
Original
Veronica Chelu, Diana Borsa, Doina Precup, and Hado van Hasselt · 2022
Later among the works it cites.
Autotelic agents with intrinsically motivated goal-conditioned reinforcement learning: a short survey
Cédric Colas, Tristan Karch, Olivier Sigaud, and Pierre-Yves Oudeyer · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de Las Casas, et al · 2022
Later among the works it cites.
Maximum entropy RL (provably) solves some robust RL problems
Benjamin Eysenbach and Sergey Levine · 2022
Later among the works it cites.
On Actions that Matter: Credit Assignment and Interpretability in Reinforcement Learning
Johan Ferret · 2022
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu · 2022
Later among the works it cites.
Deep hierarchical planning from pixels
Danijar Hafner, Kuang-Huei Lee, Ian Fischer, and Pieter Abbeel · 2022
Later among the works it cites.
Adaptive interest for emphatic reinforcement learning
Martin Klissarov, Rasool Fakoor, Jonas Mueller, Kavosh Asadi, Taesup Kim, and Alex Smola · 2022
Later among the works it cites.
On the generalization of representations in reinforcement learning
Charline Le Lan, Stephen Tu, Adam Oberman, Rishabh Agarwal, and Marc G Bellemare · 2022
Later among the works it cites.
Multi-game decision transformers
Original
Kuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee, Daniel Freeman, Winnie Xu, Sergio Guadarrama, Ian Fischer, Eric Jang, Henryk Michalewski, et al · 2022
Later among the works it cites.
Goal-conditioned reinforcement learning: Problems and solutions
Original
Minghuan Liu, Menghui Zhu, and Weinan Zhang · 2022
Later among the works it cites.
Direct advantage estimation
Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz, and Bernhard Schölkopf · 2022
Later among the works it cites.
Mastering the game of stratego with model-free multiagent reinforcement learning
Julien Perolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vincent de Boer, Paul Muller, Jerome T Connor, Neil Burch, Thomas Anthony, et al · 2022
Later among the works it cites.
A generalist agent
Original
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al · 2022
Later among the works it cites.
Learning long-term reward redistribution via randomized return decomposition
Zhizhou Ren, Ruihan Guo, Yuan Zhou, and Jian Peng · 2022
Later among the works it cites.
The phenomenon of policy churn
Original
Tom Schaul, André Barreto, John Quan, and Georg Ostrovski · 2022
Later among the works it cites.
Utility theory for sequential decision making
Mehran Shakerinava and Siamak Ravanbakhsh · 2022
Later among the works it cites.
Upside-down reinforcement learning can diverge in stochastic environments with episodic resets
Original
Miroslav Štrupl, Francesco Faccio, Dylan R Ashley, Jürgen Schmidhuber, and Rupesh Kumar Srivastava · 2022
Later among the works it cites.
How does value distribution in distributional reinforcement learning help optimization?
Original
Ke Sun, Bei Jiang, and Linglong Kong · 2022
Later among the works it cites.
Policy gradients incorporating the future
David Venuto, Elaine Lau, Doina Precup, and Ofir Nachum · 2022
Later among the works it cites.
Adversarial policies beat professional-level go ais
Original
Tony Tong Wang, Adam Gleave, Nora Belrose, Tom Tseng, Joseph Miller, Michael D Dennis, Yawen Duan, Viktor Pogrebniak, Sergey Levine, and Stuart Russell · 2022
Later among the works it cites.
Outracing champion gran turismo drivers with deep reinforcement learning
Peter R Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Thomas J Walsh, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, et al · 2022
Later among the works it cites.
Online decision transformer
Qinqing Zheng, Amy Zhang, and Aditya Grover · 2022
Later among the works it cites.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver, and Satinder Singh · 2022
Later among the works it cites.
A definition of continual reinforcement learning
David Abel, Andre Barreto, Benjamin Van Roy, Doina Precup, Hado van Hasselt, and Satinder Singh · 2023
Closest in time.
Settling the reward hypothesis
Michael Bowling, John D Martin, David Abel, and Will Dabney · 2023
Closest in time.
General intelligence requires rethinking exploration
Minqi Jiang, Tim Rocktäschel, and Edward Grefenstette · 2023
Closest in time.
Human-level atari 200x faster
Steven Kapturowski, Víctor Campos, Ray Jiang, Nemanja Rakicevic, Hado van Hasselt, Charles Blundell, and Adria Puigdomenech Badia · 2023
Closest in time.
A survey of zero-shot generalisation in deep reinforcement learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2023
Closest in time.
Quantile credit assignment
Thomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang, Mark Rowland, Theophane Weber, Clare Lyle, Audrunas Gruslys, Michal Valko, Will Dabney, Georg Ostrovski, Eric Moulines, and Remi Munos · 2023
Closest in time.
When do transformers shine in RL? decoupling memory from credit assignment
Tianwei Ni, Michel Ma, Benjamin Eysenbach, and Pierre-Luc Bacon · 2023
Closest in time.
Distributional meta-gradient reinforcement learning
Haiyan Yin, Shuicheng YAN, and Zhongwen Xu · 2023
Closest in time.
Optimizing agent behavior over long time scales by transporting value
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2041
Closest in time.