Fetching the paper…
Reading the bibliography…
Recently many algorithms were devised for reinforcement learning (RL) with function approximation.
Advantage updating
Leemon C Baird III · 1993
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2006
Earlier work this paper cites.
Probabilistic inference for solving discrete and continuous state markov decision processes
Marc Toussaint and Amos Storkey · 2006
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan Peters · 2008
Earlier work this paper cites.
General duality between optimal control and estimation
Emanuel Todorov · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altün · 2010
Earlier work this paper cites.
Variational inference for policy search in changing situations
Gerhard Neumann et al · 2011
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Earlier work this paper cites.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Stein variational gradient descent: A general purpose bayesian inference algorithm
Qiang Liu and Dilin Wang · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Earlier work this paper cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Navdeep Jaitly, Mike Schuster, Yonghui Wu, Dale Schuurmans, et al · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoffrey Roeder, and David Duvenaud · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Cited alongside, same era.
Variational inference mpc for bayesian model-based reinforcement learning
Masashi Okada and Tadahiro Taniguchi · 2019
Later among the works it cites.
Advantage-Weighted Regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
Behavior Regularized Offline Reinforcement Learning
Yifan Wu, George Tucker, and Ofir Nachum · 2019
Later among the works it cites.
Expected policy gradients for reinforcement learning
Kamil Ciosek and Shimon Whiteson · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Riashat Islam, Peter Henderson, Maziar Gomrokchi, and Doina Precup · 2017
Cited alongside, same era.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Natasha Jaques, Shixiang Gu, Dzmitry Bahdanau, José Miguel Hernández-Lobato, Richard E Turner, and Douglas Eck · 2017
Cited alongside, same era.
Action-depedent control variates for policy optimization via stein’s identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2017
Cited alongside, same era.
Improving policy gradient by exploring under-appreciated rewards
Ofir Nachum, Mohammad Norouzi, and Dale Schuurmans · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2020
Later among the works it cites.
Action and perception as divergence minimization
Danijar Hafner, Pedro A Ortega, Jimmy Ba, Thomas Parr, Karl Friston, and Nicolas Heess · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, Sarah Henderson, Alex Novikov, Sergio Gómez Colmenarejo, Serkan Cabi, Caglar Gulcehre, Tom Le Paine, Andrew Cowie, Ziyu Wang, Bilal Piot, and Nando de Freitas · 2020
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov · 2020
Later among the works it cites.
Stochastic latent actor-critic: Deep reinforcement learning with a latent variable model
Alex X. Lee, Anusha Nagabandi, Pieter Abbeel, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Making sense of reinforcement learning and probabilistic inference
Brendan O’Donoghue, Ian Osband, and Catalin Ionescu · 2020
Later among the works it cites.
Keep doing what worked: Behavioral modelling priors for offline reinforcement learning
Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, and Martin A. Riedmiller · 2020
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, Nicolas Heess, Dan Belov, Martin Riedmiller, and Matthew M. Botvinick · 2020
Later among the works it cites.
Hindsight expectation maximization for goal-conditioned reinforcement learning
Yunhao Tang and Alp Kucukelbir · 2020
Later among the works it cites.
Ziyu Wang, Alexander Novikov, Konrad Zolna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, and Nando de Freitas · 2020
Later among the works it cites.
What matters for on-policy deep actor-critic methods? a large-scale study
Marcin Andrychowicz, Anton Raichuk, Piotr Stańczyk, Manu Orsini, Sertan Girgin, Raphaël Marinier, Leonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, Sylvain Gelly, and Olivier Bachem · 2021
Closest in time.
Chainerrl: A deep reinforcement learning library
Yasuhiro Fujita, Prabhat Nagarajan, Toshiki Kataoka, and Takahiro Ishikawa · 2021
Closest in time.
Leverage the average: an analysis of kl regularization in rl
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Rémi Munos, and Matthieu Geist · 2021
Closest in time.