Fetching the paper…
Reading the bibliography…
Gradient-based optimization is the foundation of deep learning and reinforcement learning.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Compiling Fast Partial Derivatives of Functions Given by Algorithms
Bert Speelpenning · 1980
Earlier work this paper cites.
Automatic differentiation: Techniques and applications
Louis B Rall · 1981
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart and Geoffrey E Hinton · 1986
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
An analysis of actor-critic algorithms using eligibility traces: reinforcement learning with imperfect value functions
Hajime Kimura, Shigenobu Kobayashi, et al · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Joe Staines and David Barber · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
Neural variational inference and learning in belief networks
Andriy Mnih and Karol Gregor · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo J Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Tobias Glasmachers, Yi Sun, Jan Peters, and Jürgen Schmidhuber · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Brenden M Lake, Ruslan Salakhutdinov, and Joshua B Tenenbaum · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines-revised
Wojciech Zaremba and Ilya Sutskever · 2015
Variational inference for monte carlo objectives
Andriy Mnih and Danilo Rezende · 2016
Later among the works it cites.
Interpolated policy gradient: Merging on-policy and off-policy gradient estimation for deep reinforcement learning
Shixiang Gu, Tim Lillicrap, Richard E Turner, Zoubin Ghahramani, Bernhard Schölkopf, and Sergey Levine · 2017
Closest in time.
Sample-efficient policy optimization with stein control variate
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2017
Closest in time.
Reducing reparameterization gradient variance
Andrew C Miller, Nicholas J Foti, Alexander D’Amour, and Ryan P Adams · 2017
Closest in time.
Reparameterization gradients through acceptance-rejection sampling algorithms
Christian Naesseth, Francisco Ruiz, Scott Linderman, and David Blei · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Thomas Unterthiner Djork-Arné Clevert and Sepp Hochreiter · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Cited alongside, same era.
Overdispersed black-box variational inference
Francisco J.R. Ruiz, Michalis K Titsias, and David M Blei
Cited in the paper.
Control functionals for monte carlo integration
Chris J Oates, Mark Girolami, and Nicolas Chopin · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever · 2017
Closest in time.
Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J Maddison, and Jascha Sohl-Dickstein · 2017
Closest in time.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Shun Liao, Roger Grosse, and Jimmy Ba · 2017
Closest in time.
The mirage of action-dependent baselines in reinforcement learning
George Tucker, Surya Bhupatiraju, Shixiang Gu, Richard E Turner, Zoubin Ghahramani, and Sergey Levine · 2018
Closest in time.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Closest in time.