Fetching the paper…
Reading the bibliography…
We study the adaption of Soft Actor-Critic (SAC), which is considered as a state-of-the-art reinforcement learning (RL) algorithm, from continuous action space to discrete action space.
The rating of chessplayers: Past and present
Arpad E Elo and Sam Sloan · 1978
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Diversity-driven exploration strategy for deep reinforcement learning
Zhang-Wei Hong, Tzu-Yun Shann, Shih-Yang Su, Yi-Hsiang Chang, Tsu-Jui Fu, and Chun-Yi Lee · 2018
Cited alongside, same era.
Softmax deep double deterministic policy gradients
Ling Pan, Qingpeng Cai, and Longbo Huang · 2020
Later among the works it cites.
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist · 2020
Later among the works it cites.
Meta-sac: Auto-tune the entropy temperature of soft actor-critic via metagradient
Yufei Wang and Tianwei Ni · 2020
Later among the works it cites.
Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun, and Bo Cheng · 2021
Later among the works it cites.
A max-min entropy framework for reinforcement learning
Seungyul Han and Youngchul Sung · 2021
Later among the works it cites.
Denoising normalizing flow
Christian Horvat and Jean-Pascal Pfister · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Soft actor-critic for discrete action settings
Petros Christodoulou · 2019
Cited alongside, same era.
Better exploration with optimistic actor critic
Kamil Ciosek, Quan Vuong, Robert Loftin, and Katja Hofmann · 2019
Cited alongside, same era.
Improving exploration in soft-actor-critic with normalizing flows policies
Patrick Nadeem Ward, Ariella Smofsky, and Avishek Joey Bose · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2020
Cited alongside, same era.
Zhimin Hou, Kuangen Zhang, Yi Wan, Dongyu Li, Chenglong Fu, and Haoyong Yu · 2020
Cited alongside, same era.
Maxmin q-learning: Controlling the estimation bias of q-learning
Qingfeng Lan, Yangchen Pan, Alona Fyshe, and Martha White · 2020
Cited alongside, same era.
Training larger networks for deep reinforcement learning
Kei Ota, Devesh K Jha, and Asako Kanezaki · 2021
Later among the works it cites.
Target entropy annealing for discrete soft actor-critic
Yaosheng Xu, Dailin Hu, Litian Liang, Stephen McAleer, Pieter Abbeel, and Roy Fox · 2021
Later among the works it cites.
Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences
Chayan Banerjee, Zhiyong Chen, and Nasimul Noman · 2022
Closest in time.
Honor of kings arena: an environment for generalization in competitive reinforcement learning
Hua Wei, Jingxiao Chen, Xiyang Ji, Hongyang Qin, Minwen Deng, Siqin Li, Liang Wang, Weinan Zhang, Yong Yu, Liu Linc, Lanxiao Huang, Deheng Ye, QIANG FU, and Yang Wei · 2022
Closest in time.
Adaptive estimation q-learning with uncertainty and familiarity
Xiaoyu Gong, Shuai Lü, Jiayu Yu, Sheng Zhu, and Zongze Li · 2023
Closest in time.