Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) encompasses diverse paradigms, including model-based RL, policy-based RL, and value-based RL, each tailored to approximate the model, optimal policy, and optimal value function, respectively.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S · 1910
Earlier work this paper cites.
The complexity of theorem-proving procedures
Cook, S. A · 1971
Earlier work this paper cites.
Universal sequential search problems
Levin, L. A · 1973
Earlier work this paper cites.
The circuit value problem is log space complete for p
Ladner, R. E · 1975
Earlier work this paper cites.
The complexity of markov decision processes
Papadimitriou, C. H · 1987
Earlier work this paper cites.
The complexity of dynamic programming
Chow, C.-S · 1989
Earlier work this paper cites.
The computational complexity of probabilistic planning
Littman, M. L · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S · 1999
Earlier work this paper cites.
A survey of computational complexity results in systems and control
Blondel, V. D · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2001
Earlier work this paper cites.
Computational complexity: a modern approach
Arora, S · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J · 2013
Earlier work this paper cites.
Deep learning
LeCun, Y · 2015
Earlier work this paper cites.
Brockman, G · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S · 2018
Cited alongside, same era.
When to trust your model: Model-based policy optimization
Janner, M · 2019
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Liu, B · 2019
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Sun, W · 2019
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Uehara, M · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T · 2021
Later among the works it cites.
Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Zhang, Z · 2021
Later among the works it cites.
Optimistic policy optimization is provably efficient in non-stationary mdps
Zhong, H · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
Tu, S · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A · 2019
Cited alongside, same era.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Agarwal, A · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Cited alongside, same era.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S · 2022
Later among the works it cites.
A general framework for sample-efficient function approximation in reinforcement learning
Chen, Z · 2022
Later among the works it cites.
Policy learning” without”overlap: Pessimism and generalized empirical bernstein’s inequality
Jin, Y · 2022
Later among the works it cites.
Nearly optimal policy optimization with stable at any time guarantee
Wu, T · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Xiao, L · 2022
Later among the works it cites.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
Zhong, H · 2022
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Lan, G · 2023
Closest in time.
The parallelism tradeoff: Limitations of log-precision transformers
Merrill, W · 2023
Closest in time.
Rate-optimal policy optimization for linear markov decision processes
Sherman, U · 2023
Closest in time.
Bayesian design principles for frequentist sequential learning
Xu, Y · 2023
Closest in time.
Settling the sample complexity of online reinforcement learning
Zhang, Z · 2023
Closest in time.
Zhong, H · 2023
Closest in time.
On representation complexity of model-based and model-free reinforcement learning
Zhu, H · 2023
Closest in time.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Feng, G · 2024
Closest in time.