Fetching the paper…
Reading the bibliography…
Deployment efficiency is an important criterion for many real-world applications of reinforcement learning (RL).
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Ruosong Wang, Ruslan Salakhutdinov, and Lin F Yang · 2005
Earlier work this paper cites.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Adaptive educational software by applying reinforcement learning
Abdellah Bennane et al · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Concurrent pac rl
Zhaohan Guo and Emma Brunskill · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities, 2015
Joel A. Tropp · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation, 2016
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Pac reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Earlier work this paper cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
M. G. Azar, Ian Osband, and R. Munos · 2017
Earlier work this paper cites.
Exploration by random network distillation, 2018
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Earlier work this paper cites.
On oracle-efficient pac rl with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Earlier work this paper cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Is q-learning provably efficient?, 2018
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I. Jordan · 2018
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
Jinglin Chen and Nan Jiang · 2019
Cited alongside, same era.
Provably efficient rl with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Dipendra Misra, Mikael Henaff, Akshay Krishnamurthy, and John Langford · 2020
Later among the works it cites.
Accelerating reinforcement learning with learned skill priors
Karl Pertsch, Youngwoon Lee, and Joseph J Lim · 2020
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune · 2019
Cited alongside, same era.
Provably efficient maximum entropy exploration, 2019
Elad Hazan, Sham M. Kakade, Karan Singh, and Abby Van Soest · 2019
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation, 2019
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2019
Cited alongside, same era.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Reinforcement learning in healthcare: A survey
Chao Yu, Jiming Liu, and Shamim Nemati · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Later among the works it cites.
Provably efficient reward-agnostic navigation with linear value iteration
Andrea Zanette, Alessandro Lazaric, Mykel J Kochenderfer, and Emma Brunskill · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2021
Later among the works it cites.
Improved worst-case regret bounds for randomized least-squares value iteration
Priyank Agrawal, Jinglin Chen, and Nan Jiang · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning, 2021
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
A provably efficient algorithm for linear markov decision process with low switching cost, 2021
Minbo Gao, Tianle Xie, Simon S. Du, and Lin F. Yang · 2021
Later among the works it cites.
Batched neural bandits, 2021
Quanquan Gu, Amin Karbasi, Khashayar Khosravi, Vahab Mirrokni, and Dongruo Zhou · 2021
Later among the works it cites.
Online sub-sampling for reinforcement learning with general function approximation
Dingwen Kong, R. Salakhutdinov, Ruosong Wang, and Lin F. Yang · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Later among the works it cites.
Deployment-efficient reinforcement learning via model-based offline optimization
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Later among the works it cites.
Tactical optimism and pessimism for deep reinforcement learning, 2021
Ted Moskovitz, Jack Parker-Holder, Aldo Pacchiano, Michael Arbel, and Michael I. Jordan · 2021
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets, 2021
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2021
Later among the works it cites.
Linear bandits with limited adaptivity and learning distributional optimal design, 2021
Yufei Ruan, Jiaqi Yang, and Yuan Zhou · 2021
Later among the works it cites.
Made: Exploration via maximizing deviation from explored regions
Tianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian, Joseph Gonzalez, and Stuart Russell · 2021
Later among the works it cites.