Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) in episodic, factored Markov decision processes (FMDPs) is studied.
Knapsack problems: algorithms and computer implementations
Silvano Martello · 1990
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
Michael Kearns and Daphne Koller · 1999
Earlier work this paper cites.
Influence and variance of a Markov chain: Application to adaptive discretization in optimal control
Rémi Munos and Andrew Moore · 1999
Earlier work this paper cites.
Stochastic dynamic programming with factored representations
Craig Boutilier, Richard Dearden, and Moisés Goldszmidt · 2000
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
Carlos Guestrin, Daphne Koller, Ronald Parr, and Shobha Venkataraman · 2003
Earlier work this paper cites.
Inequalities for the L 1 L_{1} deviation of the empirical distribution
Tsachy Weissman, Erik Ordentlich, Gadiel Seroussi, Sergio Verdu, and Marcelo J Weinberger · 2003
Earlier work this paper cites.
Multidimensional knapsack problems
Hans Kellerer, Ulrich Pferschy, and David Pisinger · 2004
Earlier work this paper cites.
Provably efficient learning with typed parametric models
Emma Brunskill, Bethany R. Leffler, Lihong Li, Michael L. Littman, and Nicholas Roy · 2009
Earlier work this paper cites.
A unifying framework for computational reinforcement learning theory
Lihong Li · 2009
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Alexander L. Strehl, Lihong Li, and Michael L. Littman · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Bandits with knapsacks
Ashwinkumar Badanidiyuru, Robert Kleinberg, and Aleksandrs Slivkins · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Cited alongside, same era.
An efficient algorithm for contextual bandits with knapsacks, and an extension to concave objectives
Shipra Agrawal, Nikhil R Devanur, and Lihong Li · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Cited alongside, same era.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Kianté Brantley, Miroslav Dudik, Thodoris Lykouris, Sobhan Miryoosefi, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2020
Closest in time.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo R Jovanović · 2020
Closest in time.
Exploration-exploitation in constrained MDPs
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Closest in time.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Christoph Dann, Lihong Li, Wei Wei, and Emma Brunskill · 2019
Cited alongside, same era.
Q-learning with UCB exploration is sample efficient for infinite-horizon MDP
Kefan Dong, Yuanhao Wang, Xiaoyu Chen, and Liwei Wang · 2019
Cited alongside, same era.
Information-theoretic confidence bounds for reinforcement learning
Xiuyuan Lu and Benjamin Van Roy · 2019
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Lin F Yang and Mengdi Wang · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Learning in Markov decision processes under constraints
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2020
Closest in time.
Towards minimax optimal reinforcement learning in factored Markov decision processes
Yi Tian, Jian Qian, and Suvrit Sra · 2020
Closest in time.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Ruosong Wang, Russ R Salakhutdinov, and Lin Yang · 2020
Closest in time.
Ziping Xu and Ambuj Tewari · 2020
Closest in time.
Learning near optimal policies with low inherent bellman error
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer, and Emma Brunskill · 2020
Closest in time.
Constrained upper confidence reinforcement learning
Liyuan Zheng and Lillian J Ratliff · 2020
Closest in time.