Fetching the paper…
Reading the bibliography…
Modern Deep Reinforcement Learning (RL) algorithms require estimates of the maximal Q-value, which are difficult to compute in continuous domains with an infinite number of possible actions.
Limiting forms of the frequency distribution of the largest or smallest member of a sample
R. A. Fisher and L. H. C. Tippett · 1928
Earlier work this paper cites.
Introduction to the theory of statistics
Alexander McFarlane Mood · 1950
Earlier work this paper cites.
Conditional logit analysis of qualitative choice behavior
Daniel McFadden · 1972
Earlier work this paper cites.
The choice axiom after twenty years
R.Duncan Luce · 1977
Earlier work this paper cites.
Structural estimation of markov decision processes
John Rust · 1986
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1999
Earlier work this paper cites.
Estimation under linex loss function
Ahmad Parsian and SNUA Kirmani · 2002
Earlier work this paper cites.
On a new converse of jensen’s inequality
Slavko Simić · 2009
Earlier work this paper cites.
Gaussian sampling by local perturbations
George Papandreou and Alan L Yuille · 2010
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Brian D Ziebart · 2010
Earlier work this paper cites.
Entropic value-at-risk: A new coherent risk measure
A. Ahmadi-Javid · 2012
Earlier work this paper cites.
On the partition function and random maximum a-posteriori perturbations
Tamir Hazan and Tommi Jaakkola · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Infinite time horizon maximum causal entropy inverse reinforcement learning
M. Bloem and N. Bambos · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
An alternative softmax operator for reinforcement learning
Kavosh Asadi and Michael L Littman · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Gergely Neu, Anders Jonsson, and V. Gómez · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
If maxent {rl} is the answer, what is the question?, 2020
Benjamin Eysenbach and Sergey Levine · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2020
Later among the works it cites.
Leverage the average: an analysis of kl regularization in rl
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, and Matthieu Geist · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint, martin j. wainwright, cambridge university press, 2019, xvii 552 pages, £57.99, hardback isbn: 978-1-1084-9802-9
G. Alastair Young · 2020
Later among the works it cites.
Uncertainty-based offline reinforcement learning with diversified q-ensemble
Gaon An, Seungyong Moon, Jang-Hyun Kim, and Hyun Oh Song · 2021
Later among the works it cites.
Logistic q-learning
Joan Bas-Serrano, Sebastian Curi, Andreas Krause, and Gergely Neu · 2021
Later among the works it cites.
Offline rl without off-policy evaluation
David Brandfonbrener, Will Whitney, Rajesh Ranganath, and Joan Bruna · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Iq-learn: Inverse soft-q learning for imitation
Divyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song, and Stefano Ermon · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Ilya Kostrikov, Ashvin Nair, and Sergey Levine · 2021
Later among the works it cites.
Continuous control benchmark of deepmind control suite and mujoco
Qing Li · 2021
Later among the works it cites.
Implicitly regularized rl with implicit q-values
Nino Vieillard, Marcin Andrychowicz, Anton Raichuk, Olivier Pietquin, and Matthieu Geist · 2021
Later among the works it cites.