Fetching the paper…
Reading the bibliography…
Actor-critic algorithms address the dual goals of reinforcement learning (RL), policy evaluation and improvement via two separate function approximators.
Issues in using function approximation for reinforcement learning
Sebastian Thrun and Anton Schwartz · 1993
Earlier work this paper cites.
An introduction to computational learning theory
Michael J Kearns and Umesh Vazirani · 1994
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
A pac analysis of a bayesian estimator
John Shawe-Taylor and Robert C Williamson · 1997
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
PAC-Bayesian generalisation error bounds for gaussian process classification
Matthias Seeger · 2002
Earlier work this paper cites.
Pac-bayesian stochastic model selection
David A McAllester · 2003
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
PAC-Bayesian model selection for reinforcement learning
M Fard and Joelle Pineau · 2010
Earlier work this paper cites.
Double q-learning
Hado Hasselt · 2010
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mulling, and Yasemin Altun · 2010
Earlier work this paper cites.
Pac-bayesian policy evaluation for reinforcement learning
Mahdi Milani Fard, Joelle Pineau, and Csaba Szepesvári · 2012
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Maximum entropy deep inverse reinforcement learning
Markus Wulfmeier, Peter Ondruska, and Ingmar Posner · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
A primer on pac-bayesian learning
B Guedj · 2019
Later among the works it cites.
Network randomization: A simple technique for generalization in deep reinforcement learning
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Later among the works it cites.
Second order pac-bayesian bounds for the weighted majority vote
A. Masegosa, S.S. Lorenzen, C. Igel, and Y. Seldin · 2020
Later among the works it cites.
Softmax deep double deterministic policy gradients
Ling Pan, Qingpeng Cai, and Longbo Huang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Neural episodic control
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adria Puigdomenech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
PAC-Bayes Control: Synthesizing controllers that provably generalize to novel environments
Anirudha Majumdar and Maxwell Goldstein · 2018
Cited alongside, same era.
Later among the works it cites.
User-friendly introduction to pac-bayes bounds
Pierre Alquier · 2021
Later among the works it cites.
Generalization bounds for meta-learning via pac-bayes and uniform stability
Alec Farid and Anirudha Majumdar · 2021
Later among the works it cites.
Diversity actor-critic: Sample-aware entropy regularization for sample-efficient exploration
Seungyul Han and Youngchul Sung · 2021
Later among the works it cites.
Learning partially known stochastic dynamics with empirical pac bayes
Manuel Haußmann, Sebastian Gerwinn, Andreas Look, Barbara Rakitsch, and Melih Kandemir · 2021
Later among the works it cites.
Tactical optimism and pessimism for deep reinforcement learning
Ted Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel, and Michael Jordan · 2021
Later among the works it cites.
State entropy maximization with random encoders for efficient exploration
Younggyo Seo, Lili Chen, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2021
Later among the works it cites.
Probably approximately correct vision-based planning using motion primitives
Sushant Veer and Anirudha Majumdar · 2021
Later among the works it cites.
Made: Exploration via maximizing deviation from explored regions
Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian, Joseph E Gonzalez, and Stuart Russell · 2021
Later among the works it cites.
Robust Generalised Bayesian Inference for Intractable Likelihoods
Takuo Matsubara, Jeremias Knoblauch, François-Xavier Briol, and Chris J. Oates · 2022
Later among the works it cites.
Distributional Reinforcement Learning
Marc G. Bellemare, Will Dabney, and Mark Rowland · 2023
Closest in time.
Gymnasium, May 2024
Mark Towers, Jordan K Terry, Ariel Kwiatkowski, John U. Balis, Gianluca Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Jin Shen Tan, and Omar G. Younis · 2024
Closest in time.