Fetching the paper…
Reading the bibliography…
In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in general settings with function approximation.
Analysis of temporal-diffference learning with function approximation
John Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Some inequalities for information divergence and related measures of discrimination
Flemming Topsoe · 2000
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
The fixed points of off-policy td
J Kolter · 2011
Earlier work this paper cites.
Openml: Networked science in machine learning
Joaquin Vanschoren, Jan N. van Rijn, Bernd Bischl, and Luis Torgo · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Prudential life insurance assessment, 2015
Anna Montoya, BigJek14, Bull, denisedunleavy, egrad, FleetwoodHack, Imbayoh, PadraicS, Pru_Admin, tpitman, and Will Cukierski · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Zheng Wen and Benjamin Van Roy · 2017
Earlier work this paper cites.
On oracle-efficient pac rl with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Earlier work this paper cites.
Practical contextual bandits with regression oracles
Dylan Foster, Alekh Agarwal, Miroslav Dudík, Haipeng Luo, and Robert Schapire · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Earlier work this paper cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Earlier work this paper cites.
An analysis of categorical distributional reinforcement learning
Mark Rowland, Marc Bellemare, Will Dabney, Rémi Munos, and Yee Whye Teh · 2018
Cited alongside, same era.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Mohammad Sadegh Talebi and Odalric-Ambrym Maillard · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
A comparative analysis of expected and distributional reinforcement learning
Clare Lyle, Marc G Bellemare, and Pablo Samuel Castro · 2019
Cited alongside, same era.
Fully parameterized quantile function for distributional reinforcement learning
Derek Yang, Li Zhao, Zichuan Lin, Tao Qin, Jiang Bian, and Tie-Yan Liu · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Tengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro, and Alekh Agarwal · 2021
Later among the works it cites.
Learning bellman complete representations for offline policy evaluation
Jonathan Chang, Kaiwen Wang, Nathan Kallus, and Wen Sun · 2022
Later among the works it cites.
Guarantees for epsilon-greedy reinforcement learning with function approximation
Chris Dann, Yishay Mansour, Mehryar Mohri, Ayush Sekhari, and Karthik Sridharan · 2022
Later among the works it cites.
Conditionally risk-averse contextual bandits
Mónika Farsang, Paul Mineiro, and Wangda Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Autonomous navigation of stratospheric balloons using reinforcement learning
Marc G Bellemare, Salvatore Candido, Pablo Samuel Castro, Jun Gong, Marlos C Machado, Subhodeep Moitra, Sameera S Ponda, and Ziyu Wang · 2020
Cited alongside, same era.
Quantile qt-opt for risk-aware vision-based robotic grasping
Cristian Bodnar, Adrian Li, Karol Hausman, Peter Pastor, and Mrinal Kalakrishnan · 2020
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Dylan Foster and Alexander Rakhlin · 2020
Cited alongside, same era.
Tight first-and second-order regret bounds for adversarial linear bandits
Shinji Ito, Shuichi Hirahara, Tasuku Soma, and Yuichi Yoshida · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J R Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, et al · 2022
Later among the works it cites.
When is partially observable reinforcement learning not scary?
Qinghua Liu, Alan Chung, Csaba Szepesvári, and Chi Jin · 2022
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2022
Later among the works it cites.
Making linear mdps practical via contrastive representation learning
Tianjun Zhang, Tongzheng Ren, Mengjiao Yang, Joseph Gonzalez, Dale Schuurmans, and Bo Dai · 2022
Later among the works it cites.
First- and second-order bounds for adversarial linear contextual bandits
Julia Olkhovskaya, Jack Mayo, Tim van Erven, Gergely Neu, and Chen-Yu Wei · 2023
Later among the works it cites.
Information Theory: From Coding to Learning
Yury Polyanskiy and Yihong Wu · 2023
Later among the works it cites.
Spectral decomposition representation for reinforcement learning
Tongzheng Ren, Tianjun Zhang, Lisa Lee, Joseph E. Gonzalez, Dale Schuurmans, and Bo Dai · 2023
Later among the works it cites.
Distributional offline policy evaluation with predictive error guarantees
Runzhe Wu, Masatoshi Uehara, and Wen Sun · 2023
Later among the works it cites.
Settling the sample complexity of online reinforcement learning
Zihan Zhang, Yuxin Chen, Jason D Lee, and Simon S Du · 2023
Later among the works it cites.
Variance-dependent regret bounds for linear bandits and reinforcement learning: Adaptivity and computational efficiency
Heyang Zhao, Jiafan He, Dongruo Zhou, Tong Zhang, and Quanquan Gu · 2023
Later among the works it cites.
Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments
Runlong Zhou, Zhang Zihan, and Simon Shaolei Du · 2023
Later among the works it cites.