Fetching the paper…
Reading the bibliography…
The proximal policy optimization (PPO) algorithm stands as one of the most prosperous methods in the field of reinforcement learning (RL).
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Sequential batch learning in finite-action linear contextual bandits
Han, Y · 2004
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N · 2006
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Earlier work this paper cites.
On function approximation in reinforcement learning: Optimism in the face of large state spaces
Yang, Z · 2011
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J · 2013
Earlier work this paper cites.
Deep learning
LeCun, Y · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D · 2017
Earlier work this paper cites.
Soft actor-critic algorithms and applications
Haarnoja, T · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Jin, C · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Cited alongside, same era.
Superhuman ai for multiplayer poker
Brown, N · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Agarwal, A · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Cited alongside, same era.
Dynamic regret of policy optimization in non-stationary environments
Provably efficient reinforcement learning with linear function approximation under adaptivity constraints
Wang, T · 2021
Later among the works it cites.
Cautiously optimistic policy optimization and exploration with linear function approximation
Zanette, A · 2021
Later among the works it cites.
Improved variance-aware confidence sets for linear bandits and linear mixture mdp
Zhang, Z · 2021
Later among the works it cites.
Optimistic policy optimization is provably efficient in non-stationary mdps
Zhong, H · 2021
Later among the works it cites.
Vo q q l: Towards optimal regret in model-free rl with nonlinear function approximation
Agarwal, A · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fei, Y · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Cited alongside, same era.
Bandit algorithms
Lattimore, T · 2020
Cited alongside, same era.
Sample complexity of reinforcement learning using linearly combined model ensembles
Modi, A · 2020
Cited alongside, same era.
Optimistic policy optimization with bandit feedback
Shani, L · 2020
Cited alongside, same era.
Provably correct optimization and exploration with non-linear policies
Feng, F · 2021
Cited alongside, same era.
Nearly minimax optimal reinforcement learning with linear function approximation
Hu, P · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Later among the works it cites.
First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
Wagenmaker, A. J · 2022
Later among the works it cites.
Nearly optimal policy optimization with stable at any time guarantee
Wu, T · 2022
Later among the works it cites.
Computationally efficient horizon-free reinforcement learning for linear mixture mdps
Zhou, D · 2022
Later among the works it cites.
Refined regret for adversarial mdps with linear function approximation
Dai, Y · 2023
Closest in time.
Improved regret bounds for linear adversarial mdps via linear optimization
Kong, F · 2023
Closest in time.
Improved regret for efficient online reinforcement learning with linear function approximation
Sherman, U · 2023
Closest in time.