Fetching the paper…
Reading the bibliography…
Both generative adversarial networks (GAN) in unsupervised learning and actor-critic methods in reinforcement learning (RL) have gained a reputation for being difficult to optimize.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
A new method of stochastic approximation type
Boris Teodorovich Polyak · 1990
Earlier work this paper cites.
Adaptive critic designs
Danil V Prokhorov and Donald C Wunsch · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart J Russell, et al · 2000
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2001
Earlier work this paper cites.
Actor-Critic Algorirhms
VR Konda · 2002
Earlier work this paper cites.
On actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2003
Earlier work this paper cites.
Convergence rate of linear two-time-scale stochastic approximation
Vijay R Konda and John N Tsitsiklis · 2004
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Pieter Abbeel and Andrew Y Ng · 2005
Earlier work this paper cites.
An overview of bilevel optimization
Benoît Colson, Patrice Marcotte, and Gilles Savard · 2007
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Umar Syed and Robert E Schapire · 2007
Earlier work this paper cites.
Reinforcement learning in feedback control
Roland Hafner and Martin Riedmiller · 2011
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicholas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
On distinguishability criteria for estimating generative models
Ian J Goodfellow · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Towards an integration of deep learning and neuroscience
Adam Marblestone, Greg Wayne, and Konrad Kording · 2016
Closest in time.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Closest in time.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2016
Closest in time.
Jeff Donahue, Philipp Krähenbühl, and Trevor Darrell · 2016
Closest in time.
Adversarially learned inference
Vincent Dumoulin, Ishmael Belghazi, Ben Poole, Alex Lamb, Martin Arjovsky, Olivier Mastropietro, and Aaron Courville · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Autoencoding beyond pixels using a learned similarity metric
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, and Ole Winther · 2015
Cited alongside, same era.
Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, and Ian Goodfellow · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
personal communication
Thomas Degris
Cited in the paper.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Closest in time.
f-GAN: Training generative neural samplers using variational divergence minimization
Sebastian Nowozin, Botond Cseke, and Ryota Tomioka · 2016
Closest in time.
Energy-based generative adversarial networks
Junbo Zhao, Michael Mathieu, and Yann LeCun · 2016
Closest in time.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models
Chelsea Finn, Paul Christiano, Pieter Abbeel, and Sergey Levine · 2016
Closest in time.