Fetching the paper…
Reading the bibliography…
Reinforcement learning encounters many challenges when applied directly in the real world.
The optimal control of partially observable Markov processes over a finite horizon
Richard D Smallwood and Edward J Sondik · 1973
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Sample-efficient reinforcement learning of undercomplete POMDPs
Chi Jin, Sham M Kakade, Akshay Krishnamurthy, and Qinghua Liu · 2006
Earlier work this paper cites.
Autonomous inverted helicopter flight via reinforcement learning
Andrew Y Ng, Adam Coates, Mark Diel, Varun Ganapathi, Jamie Schulte, Ben Tse, Eric Berger, and Eric Liang · 2006
Earlier work this paper cites.
Evolutionary robotics
Dario Floreano, Phil Husbands, and Stefano Nolfi · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Peter L Bartlett and Ambuj Tewari · 2012
Earlier work this paper cites.
MOMDPs: a solution for modelling adaptive management problems
Iadine Chades, Josie Carwardine, Tara G Martin, Samuel Nicol, Régis Sabbadin, and Olivier Buffet · 2012
Earlier work this paper cites.
On the computational complexity of stochastic controller optimization in POMDPs
Nikos Vlassis, Michael L Littman, and David Barber · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the Eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Earlier work this paper cites.
Efficient reinforcement learning for robots using informative simulated priors
Mark Cutler and Jonathan P How · 2015
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
Paul Christiano, Zain Shah, Igor Mordatch, Jonas Schneider, Trevor Blackwell, Joshua Tobin, Pieter Abbeel, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Cad2rl: Real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2016
Earlier work this paper cites.
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Earlier work this paper cites.
Sim-to-real robot learning from pixels with progressive nets
Andrei A Rusu, Matej Večerík, Thomas Rothörl, Nicolas Heess, Razvan Pascanu, and Raia Hadsell · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture Markov decision processes
Dongruo Zhou, Quanquan Gu, and Csaba Szepesvari · 2017
Cited alongside, same era.
Using simulation and domain adaptation to improve efficiency of deep robotic grasping
Konstantinos Bousmalis, Alex Irpan, Paul Wohlhart, Yunfei Bai, Matthew Kelcey, Mrinal Kalakrishnan, Laura Downs, Julian Ibarz, Peter Pastor, Kurt Konolige, et al · 2018
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric, and Ronald Ortner · 2018
ROADS: Randomization for obstacle avoidance and driving in simulation
Samira Pouyanfar, Muneeb Saleem, Nikhil George, and Shu-Ching Chen · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zihan Zhang and Xiangyang Ji · 2019
Later among the works it cites.
PAC reinforcement learning without real-world feedback
Yuren Zhong, Aniket Anand Deshmukh, and Clayton Scott · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
PAC reinforcement learning with an imperfect model
Nan Jiang · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Sim-to-real reinforcement learning for deformable object manipulation
Jan Matas, Stephen James, and Andrew J Davison · 2018
Cited alongside, same era.
Learning dexterous in-hand manipulation
OpenAI, Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafał Józefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, Jonas Schneider, Szymon Sidor, Josh Tobin, Peter Welinder, Lilian Weng, and Wojciech Zaremba · 2018
Cited alongside, same era.
Bayesian optimization with automatic prior selection for data-efficient direct policy search
Rémi Pautrat, Konstantinos Chatzilygeroudis, and Jean-Baptiste Mouret · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Sim-to-real: Learning agile locomotion for quadruped robots
Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, and Vincent Vanhoucke · 2018
Cited alongside, same era.
Later among the works it cites.
Improved analysis of UCRL2 with empirical bernstein inequality
Ronan Fruit, Matteo Pirotta, and Alessandro Lazaric · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
Learning to generalize across long-horizon tasks from human demonstrations
Ajay Mandlekar, Danfei Xu, Roberto Martín-Martín, Silvio Savarese, and Li Fei-Fei · 2020
Later among the works it cites.
Modeling long-horizon tasks as sequential interaction landscapes
Sören Pirk, Karol Hausman, Alexander Toshev, and Mohi Khansari · 2020
Later among the works it cites.
Ruosong Wang, Ruslan Salakhutdinov, and Lin F Yang · 2020
Later among the works it cites.
Model-free reinforcement learning in infinite-horizon average-reward Markov decision processes
Chen-Yu Wei, Mehdi Jafarnia Jahromi, Haipeng Luo, Hiteshi Sharma, and Rahul Jain · 2020
Later among the works it cites.
Bellman Eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Chi Jin, Qinghua Liu, and Sobhan Miryoosefi · 2021
Closest in time.
Online sub-sampling for reinforcement learning with general function approximation
Dingwen Kong, Ruslan Salakhutdinov, Ruosong Wang, and Lin F Yang · 2021
Closest in time.
RL for latent MDPs: Regret guarantees and a lower bound
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, and Shie Mannor · 2021
Closest in time.
Haoyi Niu, Jianming Hu, Zheyu Cui, and Yi Zhang · 2021
Closest in time.
Multi-model Markov decision processes
Lauren N Steimle, David L Kaufman, and Brian T Denton · 2021
Closest in time.
Learning infinite-horizon average-reward MDPs with linear function approximation
Chen-Yu Wei, Mehdi Jafarnia Jahromi, Haipeng Luo, and Rahul Jain · 2021
Closest in time.
Sublinear regret for learning POMDPs
Yi Xiong, Ningyuan Chen, Xuefeng Gao, and Xiang Zhou · 2021
Closest in time.