Fetching the paper…
Reading the bibliography…
Entropy regularization has been extensively adopted to improve the efficiency, the stability, and the convergence of algorithms in reinforcement learning.
On the global convergence of stochastic fictitious play
Josef Hofbauer and William H Sandholm · 2002
Earlier work this paper cites.
Large-population cost-coupled lqg problems with nonuniform agents: individual-mass behavior and decentralized ϵ \epsilon -Nash equilibria
Minyi Huang, Peter E Caines, and Roland P Malhamé · 2007
Earlier work this paper cites.
Mean field games
Jean-Michel Lasry and Pierre-Louis Lions · 2007
Earlier work this paper cites.
Classes of multiagent Q-learning dynamics with epsilon-greedy exploration
Michael Wunder, Michael L Littman, and Monica Babes · 2010
Earlier work this paper cites.
Explicit solutions of some linear-quadratic mean field games
Martino Bardi · 2012
Earlier work this paper cites.
Adaptive control
Karl J Aström and Björn Wittenmark · 2013
Earlier work this paper cites.
Mean field forward-backward stochastic differential equations
René Carmona, François Delarue, et al · 2013
Earlier work this paper cites.
A probabilistic weak formulation of mean field games and applications
René Carmona, Daniel Lacker, et al · 2015
Earlier work this paper cites.
Mean field games via controlled martingale problems: existence of markovian equilibria
Daniel Lacker · 2015
Earlier work this paper cites.
Mean field games via controlled martingale problems: existence of markovian equilibria
Daniel Lacker · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Linear-quadratic mean field games
Alain Bensoussan, KCJ Sung, Sheung Chi Phillip Yam, and Siu-Pang Yung · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Earlier work this paper cites.
An information-theoretic analysis of thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
Earlier work this paper cites.
Learning in mean field games: the fictitious play
Pierre Cardaliaguet and Saeed Hadikhanloo · 2017
Cited alongside, same era.
Mean field and n-agent games for optimal investment under relative performance criteria
Daniel Lacker and Thaleia Zariphopoulou · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Cited alongside, same era.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Actor-critic provably finds Nash equilibria of linear-quadratic mean-field games
Zuyue Fu, Zhuoran Yang, Yongxin Chen, and Zhaoran Wang · 2019
Later among the works it cites.
Learning mean-field games
Xin Guo, Anran Hu, Renyuan Xu, and Junzi Zhang · 2019
Later among the works it cites.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Later among the works it cites.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Later among the works it cites.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Later among the works it cites.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Soft Q-learning with mutual-information regularization
Jordi Grau-Moya, Felix Leibfried, and Peter Vrancx · 2018
Cited alongside, same era.
A deep policy inference Q-network for multi-agent systems
Zhang-Wei Hong, Shih-Yang Su, Tzu-Yun Shann, Yi-Hsiang Chang, and Chun-Yi Lee · 2018
Cited alongside, same era.
Modeling others using oneself in multi-agent reinforcement learning
Roberta Raileanu, Emily Denton, Arthur Szlam, and Rob Fergus · 2018
Cited alongside, same era.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Cited alongside, same era.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Cited alongside, same era.
The Master Equation and the Convergence Problem in Mean Field Games:(AMS-201)
Pierre Cardaliaguet, François Delarue, Jean-Michel Lasry, and Pierre-Louis Lions · 2019
Cited alongside, same era.
Linear-quadratic mean-field reinforcement learning: Convergence of policy gradient methods
René Carmona, Mathieu Laurière, and Zongjun Tan · 2019
Cited alongside, same era.
Later among the works it cites.
Continuous-time mean-variance portfolio selection: A reinforcement learning framework
Haoran Wang and Xun Yu Zhou · 2019
Later among the works it cites.
Policy optimization provably converges to Nash equilibria in zero-sum linear quadratic games
Kaiqing Zhang, Zhuoran Yang, and Tamer Basar · 2019
Later among the works it cites.
Q-learning in regularized mean-field games
Berkay Anahtarci, Can Deha Kariksiz, and Naci Saldi · 2020
Closest in time.
Policy gradient methods for the noisy linear quadratic regulator over a finite horizon
Ben M Hambly, Renyuan Xu, and Huining Yang · 2020
Closest in time.
Exploration versus exploitation in reinforcement learning: a stochastic control approach
Haoran Wang, Thaleia Zariphopoulou, and Xunyu Zhou · 2020
Closest in time.
Weichen Wang, Jiequn Han, Zhuoran Yang, and Zhaoran Wang · 2020
Closest in time.
Policy gradient methods find the nash equilibrium in n-player general-sum linear-quadratic games
Ben M Hambly, Renyuan Xu, and Huining Yang · 2021
Closest in time.