Fetching the paper…
Reading the bibliography…
We propose training fitted Q-iteration with log-loss (FQI-log) for batch reinforcement learning (RL).
A characterization of superlinear convergence and its application to quasi-Newton methods
Dennis, J. E. and Moré, J. J · 1974
Earlier work this paper cites.
Efficient memory-based learning for robot control
Moore, A. W · 1990
Earlier work this paper cites.
Dynamic programming and optimal control
Bertsekas, D · 1995
Earlier work this paper cites.
Algorithms for sequential decision-making
Littman, M. L · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E · 1997
Earlier work this paper cites.
Some inequalities for information divergence and related measures of discrimination
Topsøe, F · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Least-squares policy iteration
Lagoudakis, M. G. and Parr, R · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
Ernst, D., Geurts, P., and Wehenkel, L · 2005
Earlier work this paper cites.
Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, M · 2005
Earlier work this paper cites.
Fitted Q-iteration in continuous action-space MDPs
Antos, A., Szepesvári, C., and Munos, R · 2007
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Reinforcement learning and simulation-based search in computer Go
Silver, D · 2009
Earlier work this paper cites.
Algorithms for reinforcement learning
Szepesvári, C · 2010
Earlier work this paper cites.
Regularization in reinforcement learning
Farahmand, A.-m · 2011
Cited alongside, same era.
Value function approximation in reinforcement learning using the Fourier basis
Konidaris, G., Osentoski, S., and Thomas, P · 2011
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
Lazaric, A., Ghavamzadeh, M., and Munos, R · 2012
Cited alongside, same era.
Statistical linear estimation with penalized estimators: An application to reinforcement learning
Pires, B. Á. and Szepesvári, C · 2012
Cited alongside, same era.
The Arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Bias no more: High-probability data-dependent regret bounds for adversarial bandits and MDPs
Lee, C.-W., Luo, H., Wei, C.-Y., and Zhang, M · 2020
Later among the works it cites.
Instance-wise minimax-optimal algorithms for logistic bandits
Abeille, M., Faury, L., and Calauzènes, C · 2021
Later among the works it cites.
The importance of pessimism in fixed-dataset policy optimization
Buckman, J., Gelada, C., and Bellemare, M. G · 2021
Later among the works it cites.
Efficient first-order contextual bandits: Prediction, allocation, and triangular discrimination
Foster, D. J. and Krishnamurthy, A · 2021
Later among the works it cites.
Offline reinforcement learning: Fundamental barriers for value function approximation
Foster, D. J., Krishnamurthy, A., Simchi-Levi, D., and Xu, Y · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
First-order regret bounds for combinatorial semi-bandits
Neu, G · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Make the minority great again: First-order regret bound for contextual bandits
Allen-Zhu, Z., Bubeck, S., and Li, Y · 2018
Cited alongside, same era.
Revisiting the Arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Is pessimism provably efficient for offline RL?
Jin, Y., Yang, Z., and Wang, Z · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Xie, T. and Jiang, N · 2021
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L. and Wang, M · 2021
Later among the works it cites.
Small-loss bounds for online learning with partial information
Lykouris, T., Sridharan, K., and Tardos, E · 2022
Later among the works it cites.
First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
Wagenmaker, A. J., Chen, Y., Simchowitz, M., Du, S., and Jamieson, K · 2022
Later among the works it cites.
Distributional reinforcement learning
Bellemare, M. G., Dabney, W., and Rowland, M · 2023
Later among the works it cites.
First- and second-order bounds for adversarial linear contextual bandits
Olkhovskaya, J., Mayo, J., van Erven, T., Neu, G., and Wei, C.-Y · 2023
Later among the works it cites.
The benefits of being distributional: Small-loss bounds for reinforcement learning
Wang, K., Zhou, K., Wu, R., Kallus, N., and Sun, W · 2023
Later among the works it cites.
Stop regressing: Training value functions via classification for scalable deep RL
Farebrother, J., Orbay, J., Vuong, Q., Taïga, A. A., Chebotar, Y., Xiao, T., Irpan, A., Levine, S., Castro, P. S., Faust, A., Kumar, A., and Agarwal, R · 2024
Closest in time.
Exploration via linearly perturbed loss minimisation
Janz, D., Liu, S., Ayoub, A., and Szepesvári, C · 2024
Closest in time.
More benefits of being distributional: Second-order bounds for reinforcement learning
Wang, K., Oertell, O., Agarwal, A., Kallus, N., and Sun, W · 2024
Closest in time.