Fetching the paper…
Reading the bibliography…
Q-learning played a foundational role in the field reinforcement learning (RL).
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Learning from Delayed Rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
The convergence of TD( λ \lambda ) for general λ \lambda
Peter Dayan · 1992
Earlier work this paper cites.
Technical note: q -learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Incremental multi-step q-learning
Jing Peng and Ronald J. Williams · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Convergence of reinforcement learning with general function approximators
Vassilis A. Papavassiliou and Stuart Russell · 1999
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
General state space Markov chains and MCMC algorithms
Gareth O. Roberts and Jeffrey S. Rosenthal · 2004
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
Matthew W. Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Nikola Momchev, Danila Sinopalnikov, Piotr Stańczyk, Sabela Ramos, Anton Raichuk, Damien Vincent, Léonard Hussenot, Robert Dadashi, Gabriel Dulac-Arnold, Manu Orsini, Alexis Jacq, Johan Ferret, Nino Vieillard, Seyed Kamyar Seyed Ghasemipour, Sertan Girgin, Olivier Pietquin, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Abe Friesen, Ruba Haroun, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2006
Earlier work this paper cites.
Stochastic Approximation: A Dynamical Systems Viewpoint
Vivek Borkar · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Stochastic approximation: a survey
Harold J. Kushner · 2010
Earlier work this paper cites.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S. Sutton · 2010
Earlier work this paper cites.
The fixed points of off-policy td
J. Kolter · 2011
Earlier work this paper cites.
A simpler approach to obtaining an o(1/t) convergence rate for the projected stochastic subgradient method
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Remi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas · 2016
Cited alongside, same era.
Finite sample analysis for TD(0) with linear function approximation
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
The nethack learning environment
Heinrich Küttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
On the global optimality of model-agnostic meta-learning: reinforcement learning and supervised learning
Lingxiao Wang, Qi Cai, Zhuoyan Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Randomized ensembled double q-learning: Learning fast without a model
Xinyue Chen, Che Wang, Zijian Zhou, and Keith W. Ross · 2021
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
Danijar Hafner · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Cited alongside, same era.
Dopamine: A Research Framework for Deep Reinforcement Learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Soft Actor-Critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Cited alongside, same era.
Hengyuan Hu, Adam Lerer, Brandon Cui, Luis Pineda, Noam Brown, and Jakob Foerster · 2021
Later among the works it cites.
Revisiting peng’s q(lambda) for modern reinforcement learning
Tadashi Kozuno, Yunhao Tang, Mark Rowland, Remi Munos, Steven Kapturowski, Will Dabney, Michal Valko, and David Abel · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning, 2021
Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, and Gavriel State · 2021
Later among the works it cites.
Breaking the deadly triad with a target network
Shangtong Zhang, Hengshuai Yao, and Shimon Whiteson · 2021
Later among the works it cites.
The 37 implementation details of proximal policy optimization
Shengyi Huang, Rousslan Fernand Julien Dossa, Antonin Raffin, Anssi Kanervisto, and Weixun Wang · 2022
Later among the works it cites.
gymnax: A JAX-based reinforcement learning environment library, 2022
Robert Tjarko Lange · 2022
Later among the works it cites.
Discovered policy optimisation
Chris Lu, Jakub Kuba, Alistair Letcher, Luke Metz, Christian Schroeder de Witt, and Jakob Foerster · 2022
Later among the works it cites.
EnvPool: A highly parallel reinforcement learning environment execution engine
Jiayi Weng, Min Lin, Shengyi Huang, Bo Liu, Denys Makoviichuk, Viktor Makoviychuk, Zichen Liu, Yufan Song, Ting Luo, Yukun Jiang, Zhongwen Xu, and Shuicheng Yan · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Atari-5: Distilling the arcade learning environment down to five games
Matthew Aitchison, Penny Sweetser, and Marcus Hutter · 2023
Later among the works it cites.
Why target networks stabilise temporal difference methods
Mattie Fellows, Matthew Smith, and Shimon Whiteson · 2023
Later among the works it cites.
Maniskill2: A unified benchmark for generalizable manipulation skills
Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen, and Hao Su · 2023
Later among the works it cites.
Parallel q q -learning: Scaling off-policy reinforcement learning under massively parallel simulation
Zechu Li, Tao Chen, Zhang-Wei Hong, Anurag Ajay, and Pulkit Agrawal · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Later among the works it cites.
Jaxmarl: Multi-agent rl environments in jax
Alexander Rutherford, Benjamin Ellis, Matteo Gallici, Jonathan Cook, Andrei Lupu, Gardar Ingvarsson, Timon Willi, Akbir Khan, Christian Schroeder de Witt, Alexandra Souly, et al · 2023
Later among the works it cites.
Flashbax: Streamlining experience replay buffers for reinforcement learning with jax, 2023
Edan Toledo, Laurence Midgley, Donal Byrne, Callum Rhys Tilbury, Matthew Macfarlane, Cyprien Courtot, and Alexandre Laterre · 2023
Later among the works it cites.
Gymnasium, March 2023
Mark Towers, Jordan K. Terry, Ariel Kwiatkowski, John U. Balis, Gianluca de Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Tan Jin Shen, and Omar G. Younis · 2023
Later among the works it cites.
Understanding, predicting and better resolving q-value divergence in offline-rl
Yang Yue, Rui Lu, Bingyi Kang, Shiji Song, and Gao Huang · 2023
Later among the works it cites.
Crossq: Batch normalization in deep reinforcement learning for greater sample efficiency and simplicity
Aditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus, Artemij Amiranashvili, Thomas Brox, and Jan Peters · 2024
Closest in time.
Jumanji: a diverse suite of scalable reinforcement learning environments in jax, 2024
Clément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana, Sasha Abramowitz, Paul Duckworth, Vincent Coyette, Laurence I. Midgley, Elshadai Tegegn, Tristan Kalloniatis, Omayma Mahjoub, Matthew Macfarlane, Andries P. Smit, Nathan Grinsztajn, Raphael Boige, Cemlyn N. Waters, Mohamed A. Mimouni, Ulrich A. Mbou Sob, Ruan de Kock, Siddarth Singh, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, and Alexandre Laterre · 2024
Closest in time.
SMACv2: An improved benchmark for cooperative multi-agent reinforcement learning
Benjamin Ellis, Jonathan Cook, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob Foerster, and Shimon Whiteson · 2024
Closest in time.
Disentangling the causes of plasticity loss in neural networks
Clare Lyle, Zeyu Zheng, Khimya Khetarpal, Hado van Hasselt, Razvan Pascanu, James Martens, and Will Dabney · 2024
Closest in time.
Bigger, regularized, optimistic: scaling for compute and sample-efficient continuous control, 2024
Michal Nauman, Mateusz Ostaszewski, Krzysztof Jankowski, Piotr Miłoś, and Marek Cygan · 2024
Closest in time.