Fetching the paper…
Reading the bibliography…
The success of deep reinforcement learning (DRL) lies in its ability to learn a representation that is well-suited for the exploration and exploitation task.
Weighted sums of certain dependent random variables
Kazuoki Azuma · 1967
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
Satinder P Singh Tommi Jaakkola Michael and I Jordan · 1995
Earlier work this paper cites.
Model minimization in markov decision processes
Thomas Dean and Robert Givan · 1997
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Model minimization in hierarchical reinforcement learning
Balaraman Ravindran and Andrew G Barto · 2002
Earlier work this paper cites.
Local rademacher complexities
Peter L Bartlett, Olivier Bousquet, Shahar Mendelson, et al · 2005
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Earlier work this paper cites.
Dylan J Foster, Alexander Rakhlin, David Simchi-Levi, and Yunzong Xu · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Adaptive bandits: Towards the best history-dependent strategy
Maillard Odalric and Rémi Munos · 2011
Earlier work this paper cites.
Regret bound balancing and elimination for model selection in bandits and rl
Aldo Pacchiano, Christoph Dann, Claudio Gentile, and Peter Bartlett · 2012
Earlier work this paper cites.
User-friendly tail bounds for sums of random matrices
Joel A Tropp · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Near optimal behavior via approximate state abstraction
David Abel, David Hershkowitz, and Michael Littman · 2016
Earlier work this paper cites.
Corralling a band of bandit algorithms
Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, and Robert E Schapire · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Integrating state representation learning into deep reinforcement learning
Tim de Bruin, Jens Kober, Karl Tuyls, and Robert Babuška · 2018
Earlier work this paper cites.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Cited alongside, same era.
Model selection for contextual bandits
Dylan Foster, Akshay Krishnamurthy, and Haipeng Luo · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin Jamieson · 2019
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Cited alongside, same era.
Learning near optimal policies with low inherent bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Simon Du, Sham Kakade, Jason Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, and Ruosong Wang · 2021
Closest in time.
Offline reinforcement learning: Fundamental barriers for value function approximation
Dylan J Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu · 2021
Closest in time.
Problem-complexity adaptive model selection for stochastic linear bandits
Avishek Ghosh, Abishek Sankararaman, and Ramchandran Kannan · 2021
Closest in time.
Logarithmic regret for reinforcement learning with linear function approximation
Jiafan He, Dongruo Zhou, and Quanquan Gu · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yasin Abbasi-Yadkori, Aldo Pacchiano, and My Phan · 2020
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham M Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2020
Cited alongside, same era.
Osom: A simultaneously optimal algorithm for multi-armed and linear contextual bandits
Niladri Chatterji, Vidya Muthukumar, and Peter Bartlett · 2020
Cited alongside, same era.
Is a good representation sufficient for sample efficient reinforcement learning?
Simon S Du, Sham M Kakade, Ruosong Wang, and Lin F Yang · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Zeyu Jia, Lin Yang, Csaba Szepesvari, and Mengdi Wang · 2020
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Ying Jin, Zhuoran Yang, and Zhaoran Wang · 2021
Closest in time.
Variance-aware off-policy evaluation with linear function approximation
Yifei Min, Tianhao Wang, Dongruo Zhou, and Quanquan Gu · 2021
Closest in time.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Closest in time.
Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Closest in time.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Closest in time.
Learning infinite-horizon average-reward mdps with linear function approximation
Chen-Yu Wei, Mehdi Jafarnia Jahromi, Haipeng Luo, and Rahul Jain · 2021
Closest in time.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Closest in time.
Q-learning with logarithmic regret
Kunhe Yang, Lin Yang, and Simon Du · 2021
Closest in time.
Near-optimal offline reinforcement learning via double variance reduction
Ming Yin, Yu Bai, and Yu-Xiang Wang · 2021
Closest in time.
Contrastive ucb: Provably efficient contrastive self-supervised learning in online reinforcement learning
Shuang Qiu, Lingxiao Wang, Chenjia Bai, Zhuoran Yang, and Zhaoran Wang · 2022
Closest in time.
On gap-dependent bounds for offline reinforcement learning
Xinqi Wang, Qiwen Cui, and Simon S Du · 2022
Closest in time.
Near-optimal offline reinforcement learning with linear representation: Leveraging variance information with pessimism
Ming Yin, Yaqi Duan, Mengdi Wang, and Yu-Xiang Wang · 2022
Closest in time.
Making linear mdps practical via contrastive representation learning
Tianjun Zhang, Tongzheng Ren, Mengjiao Yang, Joseph Gonzalez, Dale Schuurmans, and Bo Dai · 2022
Closest in time.