Fetching the paper…
Reading the bibliography…
Learning a transition model via Maximum Likelihood Estimation (MLE) followed by planning inside the learned model is perhaps the most standard and simplest Model-based Reinforcement Learning (RL) framework.
Central limit theorems for empirical measures
Richard M Dudley · 1978
Earlier work this paper cites.
Task-level robot learning: Juggling a tennis ball more accurately
Eric W Aboaf, Steven Mark Drucker, and Christopher G Atkeson · 1989
Earlier work this paper cites.
Empirical Processes in M-estimation , volume 6
Sara A Geer · 2000
Earlier work this paper cites.
From ε \varepsilon -entropy to kl-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Improved second-order bounds for prediction with expert advice
Nicolo Cesa-Bianchi, Yishay Mansour, and Gilles Stoltz · 2007
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Marc Deisenroth, Carl Rasmussen, and Dieter Fox · 2011
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and J Andrew Bagnell · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Improved learning of dynamics models for control
Arun Venkatraman, Roberto Capobianco, Lerrel Pinto, Martial Hebert, Daniele Nardi, and J Andrew Bagnell · 2017
Earlier work this paper cites.
Information theoretic mpc for model-based reinforcement learning
Grady Williams, Nolan Wagener, Brian Goldfain, Paul Drews, James M Rehg, Byron Boots, and Evangelos A Theodorou · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Earlier work this paper cites.
Model-based reinforcement learning for atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Earlier work this paper cites.
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht · 2019
Earlier work this paper cites.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Earlier work this paper cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Tight first-and second-order regret bounds for adversarial linear bandits
Shinji Ito, Shuichi Hirahara, Tasuku Soma, and Yuichi Yoshida · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Ruosong Wang, Simon S Du, Lin F Yang, and Sham M Kakade · 2020
Cited alongside, same era.
Pac reinforcement learning for predictive state representations
Wenhao Zhan, Masatoshi Uehara, Wen Sun, and Jason D Lee · 2022
Later among the works it cites.
Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Zihan Zhang, Xiangyang Ji, and Simon Du · 2022
Later among the works it cites.
Gec: A unified framework for interactive decision making in mdp, pomdp, and beyond
Han Zhong, Wei Xiong, Sirui Zheng, Liwei Wang, Zhaoran Wang, Zhuoran Yang, and Tong Zhang · 2022
Later among the works it cites.
Computationally efficient horizon-free reinforcement learning for linear mixture mdps
Dongruo Zhou and Quanquan Gu · 2022
Later among the works it cites.
Nearly minimax optimal regret for learning linear mixture stochastic shortest path
Qiwei Di, Jiafan He, Dongruo Zhou, and Quanquan Gu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
Dylan J Foster and Akshay Krishnamurthy · 2021
Cited alongside, same era.
The statistical complexity of interactive decision making
Dylan J Foster, Sham M Kakade, Jian Qian, and Alexander Rakhlin · 2021
Cited alongside, same era.
Nearly horizon-free offline reinforcement learning
Tongzheng Ren, Jialian Li, Bo Dai, Simon S Du, and Sujay Sanghavi · 2021
Cited alongside, same era.
Pc-mlp: Model-based reinforcement learning with policy cover guided exploration
Yuda Song and Wen Sun · 2021
Cited alongside, same era.
Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret
Jean Tarbouriech, Runlong Zhou, Simon S Du, Matteo Pirotta, Michal Valko, and Alessandro Lazaric · 2021
Cited alongside, same era.
Pessimistic model-based offline reinforcement learning under partial coverage
Masatoshi Uehara and Wen Sun · 2021
Cited alongside, same era.
Representation learning for online and offline rl in low-rank mdps
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2021
Cited alongside, same era.
Optimistic mle: A generic model-based algorithm for partially observable sequential decision making
Qinghua Liu, Praneeth Netrapalli, Csaba Szepesvari, and Chi Jin · 2023
Later among the works it cites.
The benefits of being distributional: Small-loss bounds for reinforcement learning
Kaiwen Wang, Kevin Zhou, Runzhe Wu, Nathan Kallus, and Wen Sun · 2023
Later among the works it cites.
Learning interactive real-world simulators
Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel · 2023
Later among the works it cites.
Optimal horizon-free reward-free exploration for linear mixture mdps
Junkai Zhang, Weitong Zhang, and Quanquan Gu · 2023
Later among the works it cites.
Runlong Zhou, Zhang Zihan, and Simon Shaolei Du · 2023
Later among the works it cites.
Is behavior cloning all you need? understanding horizon in imitation learning
Dylan J Foster, Adam Block, and Dipendra Misra · 2024
Closest in time.
Horizon-free and instance-dependent regret bounds for reinforcement learning with general function approximation
Jiayi Huang, Han Zhong, Liwei Wang, and Lin Yang · 2024
Closest in time.
First-and second-order bounds for adversarial linear contextual bandits
Julia Olkhovskaya, Jack Mayo, Tim van Erven, Gergely Neu, and Chen-Yu Wei · 2024
Closest in time.
More benefits of being distributional: Second-order bounds for reinforcement learning
Kaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus, and Wen Sun · 2024
Closest in time.
Corruption-robust offline reinforcement learning with general function approximation
Chenlu Ye, Rui Yang, Quanquan Gu, and Tong Zhang · 2024
Closest in time.