Fetching the paper…
Reading the bibliography…
While there has been substantial success for solving continuous control with actor-critic methods, simpler critic-only methods such as Q-learning find limited application in the associated high-dimensional action spaces.
On the “bang-bang” control problem
Richard Bellman, Irving Glicksberg, and Oliver Gross · 1956
Earlier work this paper cites.
The ‘bang-bang’principle
Joseph P LaSalle · 1960
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Distributed value functions
Jeff G Schneider, Weng-Keen Wong, Andrew W Moore, and Martin A Riedmiller · 1999
Earlier work this paper cites.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Earlier work this paper cites.
Coordinated reinforcement learning
Carlos Guestrin, Michail Lagoudakis, and Ronald Parr · 2002
Earlier work this paper cites.
Q-decomposition for reinforcement learning agents
Stuart J Russell and Andrew Zimdars · 2003
Earlier work this paper cites.
Decentralized reinforcement learning control of a robotic manipulator
Lucian Busoniu, Bart De Schutter, and Robert Babuska · 2006
Earlier work this paper cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2006
Earlier work this paper cites.
Collaborative multiagent reinforcement learning by payoff propagation
Jelle R Kok and Nikos Vlassis · 2006
Earlier work this paper cites.
Lenient learners in cooperative multiagent systems
Liviu Panait, Keith Sullivan, and Sean Luke · 2006
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Laëtitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2007
Earlier work this paper cites.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Xiangxiang Chu and Hangjun Ye · 2017
Earlier work this paper cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Earlier work this paper cites.
Cooperative multi-agent control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel Kochenderfer · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Cited alongside, same era.
Discrete sequential prediction of continuous actions for deep rl
Luke Metz, Julian Ibarz, Navdeep Jaitly, and James Davidson · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Later among the works it cites.
Continuous-discrete reinforcement learning for hybrid control in robotics
Michael Neunert, Abbas Abdolmaleki, Markus Wulfmeier, Thomas Lampe, Tobias Springenberg, Roland Hafner, Francesco Romano, Jonas Buchli, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
Curl: Contrastive unsupervised representations for reinforcement learning
Aravind Srinivas, Michael Laskin, and Pieter Abbeel · 2020
Later among the works it cites.
Discretizing continuous action space for on-policy optimization
Yunhao Tang and Shipra Agrawal · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sahil Sharma, Aravind Suresh, Rahul Ramesh, and Balaraman Ravindran · 2017
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Cited alongside, same era.
Hybrid reward architecture for reinforcement learning
Harm Van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva Tb, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Cited alongside, same era.
dm_control: Software and tasks for continuous control
Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, Nicolas Heess, and Yuval Tassa · 2020
Later among the works it cites.
Q-learning in enormous action spaces via amortized approximate maximization
Tom Van de Wiele, David Warde-Farley, Andriy Mnih, and Volodymyr Mnih · 2020
Later among the works it cites.
Dop: Off-policy multi-agent decomposed policy gradients
Yihan Wang, Beining Han, Tonghan Wang, Heng Dong, and Chongjie Zhang · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Cooperative prioritized sweeping
Eugenio Bargiacchi, Timothy Verstraeten, and Diederik M Roijers · 2021
Later among the works it cites.
Scaling multi-agent reinforcement learning with selective parameter sharing
Filippos Christianos, Georgios Papoudakis, Muhammad A Rahman, and Stefano V Albrecht · 2021
Later among the works it cites.
Isaac gym: High performance gpu-based physics simulation for robot learning
Viktor Makoviychuk, Lukasz Wawrzyniak, Yunrong Guo, Michelle Lu, Kier Storey, Miles Macklin, David Hoeller, Nikita Rudin, Arthur Allshire, Ankur Handa, et al · 2021
Later among the works it cites.
Facmac: Factored multi-agent centralised policy gradients
Bei Peng, Tabish Rashid, Christian Schroeder de Witt, Pierre-Alexandre Kamienny, Philip Torr, Wendelin Böhmer, and Shimon Whiteson · 2021
Later among the works it cites.
Is bang-bang control all you need? solving continuous control with bernoulli policies
Tim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato, Martin Riedmiller, Markus Wulfmeier, and Daniela Rus · 2021
Later among the works it cites.
Value-decomposition multi-agent actor-critics
Jianyu Su, Stephen Adams, and Peter A Beling · 2021
Later among the works it cites.
On structural and temporal credit assignment in reinforcement learning
Arash Tavakoli · 2021
Later among the works it cites.
Learning to represent action values as a hypergraph on the action vertices
Arash Tavakoli, Mehdi Fatemi, and Petar Kormushev · 2021
Later among the works it cites.
Representation matters: Improving perception and exploration for robotics
Markus Wulfmeier, Arunkumar Byravan, Tim Hertweck, Irina Higgins, Ankush Gupta, Tejas D. Kulkarni, Malcolm Reynolds, Denis Teplyashin, Roland Hafner, Thomas Lampe, and Martin A. Riedmiller · 2021
Later among the works it cites.
Accelerated policy learning with parallel differentiable simulation
Jie Xu, Viktor Makoviychuk, Yashraj Narang, Fabio Ramos, Wojciech Matusik, Animesh Garg, and Miles Macklin · 2021
Later among the works it cites.
Mastering visual continuous control: Improved data-augmented reinforcement learning
Denis Yarats, Rob Fergus, Alessandro Lazaric, and Lerrel Pinto · 2021
Later among the works it cites.
Temporal difference learning for model predictive control
Nicklas Hansen, Xiaolong Wang, and Hao Su · 2022
Closest in time.
Learning to walk in minutes using massively parallel deep reinforcement learning
Nikita Rudin, David Hoeller, Philipp Reist, and Marco Hutter · 2022
Closest in time.