Fetching the paper…
Reading the bibliography…
Concave Utility Reinforcement Learning (CURL) extends RL from linear to concave utilities in the occupancy measure induced by the agent's policy.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
An algorithm for quadratic programming
Marguerite Frank and Philip Wolfe · 1956
Earlier work this paper cites.
Solving asymmetric variational inequality problems and systems of equations with generalized nonlinear programming algorithms
Janice H Hammond · 1984
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
Large population stochastic dynamic games: closed-loop mckean-vlasov systems and the nash certainty equivalence principle
Minyi Huang, Roland P Malhamé, Peter E Caines, et al · 2006
Earlier work this paper cites.
Mean field games
Jean-Michel Lasry and Pierre-Louis Lions · 2007
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Extended deterministic mean-field games
Diogo A Gomes and Vardan K Voskanyan · 2016
Earlier work this paper cites.
On frank-wolfe and equilibrium computation
Jacob Abernethy and Jun-Kun Wang · 2017
Earlier work this paper cites.
Learning in mean field games: the fictitious play
Pierre Cardaliaguet and Saeed Hadikhanloo · 2017
Earlier work this paper cites.
Frank-wolfe algorithms for saddle point problems
Gauthier Gidel, Tony Jebara, and Simon Lacoste-Julien · 2017
Earlier work this paper cites.
Numerical methods for mean-field-type optimal control problems
Laurent Pfeiffer · 2017
Earlier work this paper cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Probabilistic Theory of Mean Field Games with Applications I-II
René Carmona, François Delarue, et al · 2018
Cited alongside, same era.
Decentralised learning in systems with many, many strategic agents
David Mguni, Joel Jennings, and Enrique Munoz de Cote · 2018
Cited alongside, same era.
Reinforcement learning for joint optimization of multiple rewards
Mridul Agarwal and Vaneet Aggarwal · 2019
Cited alongside, same era.
Fitted q-learning in mean-field games
Berkay Anahtarcı, Can Deha Karıksız, and Naci Saldi · 2019
Cited alongside, same era.
Connecting gans, mfgs, and ot
Haoyang Cao, Xin Guo, and Mathieu Laurière · 2020
Later among the works it cites.
On the convergence of model free learning in mean field games
Romuald Elie, Julien Pérolat, Mathieu Laurière, Matthieu Geist, and Olivier Pietquin · 2020
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Fictitious play for mean field games: Continuous time analysis and applications
Sarah Perrin, Julien Pérolat, Mathieu Laurière, Matthieu Geist, Romuald Elie, and Olivier Pietquin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regret minimization for reinforcement learning with vectorial feedback and complex objectives
Wang Chi Cheung · 2019
Cited alongside, same era.
A theory of regularized markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Cited alongside, same era.
Learning mean-field games
Xin Guo, Anran Hu, Renyuan Xu, and Junzi Zhang · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration
Elad Hazan, Sham Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Efficient exploration via state marginal matching
Lisa Lee, Benjamin Eysenbach, Emilio Parisotto, Eric Xing, Sergey Levine, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
Reinforcement learning in stationary mean-field games
Jayakumar Subramanian and Aditya Mahajan · 2019
Cited alongside, same era.
A mean-field optimal control formulation of deep learning
E Weinan, Jiequn Han, and Qianxiao Li · 2019
Cited alongside, same era.
Ahmed Touati, Amy Zhang, Joelle Pineau, and Pascal Vincent · 2020
Later among the works it cites.
Provable fictitious play for general mean-field games
Qiaomin Xie, Zhuoran Yang, Zhaoran Wang, and Andreea Minca · 2020
Later among the works it cites.
Variational policy gradient method for reinforcement learning with general utilities
Junyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvari, and Mengdi Wang · 2020
Later among the works it cites.
Approximately solving mean field games via entropy-regularized deep reinforcement learning
Kai Cui and Heinz Koeppl · 2021
Closest in time.
Deep learning and mean-field games: A stochastic optimal control perspective
Luca Di Persio and Matteo Garbelli · 2021
Closest in time.
Normalizing flows for probabilistic modeling and inference
George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan · 2021
Closest in time.
Reward is enough for convex mdps
Tom Zahavy, Brendan O’Donoghue, Guillaume Desjardins, and Satinder Singh · 2021
Closest in time.
Scaling up mean field games with online mirror descent
Julien Perolat, Sarah Perrin, Romuald Elie, Mathieu Laurière, Georgios Piliouras, Matthieu Geist, Karl Tuyls, and Olivier Pietquin · 2022
Closest in time.