Fetching the paper…
Reading the bibliography…
Reinforcement learning typically assumes that agents observe feedback for their actions immediately, but in many real-world applications (like recommendation systems) feedback is observed in delay.
Introduction to online convex optimization
Hazan, E. 2019 · 1909
Earlier work this paper cites.
Frequentist regret bounds for randomized least-squares value iteration
Zanette, A.; Brandfonbrener, D.; Brunskill, E.; Pirotta, M.; and Lazaric, A. 2020a · 1964
Earlier work this paper cites.
On delayed prediction of individual sequences
Weinberger, M. J.; and Ordentlich, E. 2002 · 1976
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A.; and Teboulle, M. 2003 · 2003
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
Katsikopoulos, K. V.; and Engelbrecht, S. E. 2003 · 2003
Earlier work this paper cites.
Delay-aware multi-agent reinforcement learning
Chen, B.; Xu, M.; Liu, Z.; Li, L.; and Zhao, D. 2020 · 2005
Earlier work this paper cites.
Online Markov decision processes
Even-Dar, E.; Kakade, S. M.; and Mansour, Y. 2009 · 2009
Earlier work this paper cites.
Empirical Bernstein Bounds and Sample Variance Penalization
Maurer, A.; and Pontil, M. 2009 · 2009
Earlier work this paper cites.
Learning and planning in environments with delayed feedback
Walsh, T. J.; Nouri, A.; Li, L.; and Littman, M. L. 2009 · 2009
Earlier work this paper cites.
Adapting to Delays and Data in Adversarial Multi-Armed Bandits
György, A.; and Joulani, P. 2020 · 2010
Earlier work this paper cites.
Near-optimal Regret Bounds for Reinforcement Learning
Jaksch, T.; Ortner, R.; and Auer, P. 2010 · 2010
Earlier work this paper cites.
The Online Loop-free Stochastic Shortest-Path Problem
Neu, G.; György, A.; and Szepesvári, C. 2010 · 2010
Earlier work this paper cites.
Control delay in reinforcement learning for real-time dynamic systems: a memoryless approach
Schuitema, E.; Buşoniu, L.; Babuška, R.; and Jonker, P. 2010 · 2010
Earlier work this paper cites.
Online learning for QoE-based video streaming to mobile receivers
Changuel, N.; Sayadi, B.; and Kieffer, M. 2012 · 2012
Earlier work this paper cites.
The adversarial stochastic shortest path problem with unknown transition probabilities
Neu, G.; György, A.; and Szepesvári, C. 2012 · 2012
Earlier work this paper cites.
Online learning under delayed feedback
Joulani, P.; Gyorgy, A.; and Szepesvári, C. 2013 · 2013
Earlier work this paper cites.
Online Learning in Markovian Decision Processes
Zimin, A. 2013 · 2013
Cited alongside, same era.
Online learning in episodic Markovian decision processes by relative entropy policy search
Zimin, A.; and Neu, G. 2013 · 2013
Cited alongside, same era.
Online Markov Decision Processes Under Bandit Feedback
Neu, G.; György, A.; Szepesvári, C.; and Antos, A. 2014 · 2014
Cited alongside, same era.
Online learning with adversarial delays
Quanrud, K.; and Khashabi, D. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Delay and cooperation in nonstochastic bandits
Cesa-Bianchi, N.; Gentile, C.; Mansour, Y.; and Minora, A. 2016 · 2016
Cited alongside, same era.
Problem dependent reinforcement learning bounds which can identify bandit structure in mdps
Zanette, A.; and Brunskill, E. 2018 · 2018
Later among the works it cites.
Online exp3 learning in adversarial bandits with delayed feedback
Bistritz, I.; Zhou, Z.; Chen, X.; Bambos, N.; and Blanchet, J. 2019 · 2019
Later among the works it cites.
Nonstochastic multiarmed bandits with unrestricted delays
Thune, T. S.; Cesa-Bianchi, N.; and Seldin, Y. 2019 · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L.; and Wang, M. 2019 · 2019
Later among the works it cites.
Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds
Zanette, A.; and Brunskill, E. 2019 · 2019
Later among the works it cites.
Learning in generalized linear contextual bandits with stochastic delays
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax regret bounds for reinforcement learning
Azar, M. G.; Osband, I.; and Munos, R. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Stochastic Bandit Models for Delayed Conversions
Vernade, C.; Cappé, O.; and Perchet, V. 2017 · 2017
Cited alongside, same era.
Nonstochastic bandits with composite anonymous feedback
Cesa-Bianchi, N.; Gentile, C.; and Mansour, Y. 2018 · 2018
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C.; Allen-Zhu, Z.; Bubeck, S.; and Jordan, M. I. 2018 · 2018
Cited alongside, same era.
Zhou, Z.; Xu, R.; and Blanchet, J. 2019 · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q.; Yang, Z.; Jin, C.; and Wang, Z. 2020 · 2020
Closest in time.
Stochastic bandits with arm-dependent delays
Gael, M. A.; Vernade, C.; Carpentier, A.; and Valko, M. 2020 · 2020
Closest in time.
A modular analysis of adaptive (non-) convex optimization: Optimism, composite objectives, variance reduction, and variational bounds
Joulani, P.; György, A.; and Szepesvári, C. 2020 · 2020
Closest in time.
Near-optimal Regret Bounds for Stochastic Shortest Path
Rosenberg, A.; Cohen, A.; Mansour, Y.; and Kaplan, H. 2020 · 2020
Closest in time.
Optimistic Policy Optimization with Bandit Feedback
Shani, L.; Efroni, Y.; Rosenberg, A.; and Mannor, S. 2020 · 2020
Closest in time.
An optimal algorithm for adversarial bandits with arbitrary delays
Zimmert, J.; and Seldin, Y. 2020 · 2020
Closest in time.
Nearly Optimal Regret for Learning Adversarial MDPs with Linear Function Approximation
He, J.; Zhou, D.; and Gu, Q. 2021 · 2021
Closest in time.
Stochastic Multi-Armed Bandits with Unrestricted Delay Distributions
Lancewicki, T.; Segal, S.; Koren, T.; and Mansour, Y. 2021 · 2021
Closest in time.
Impact of communication delays on secondary frequency control in an islanded microgrid
Liu, S.; Wang, X.; and Liu, P. X. 2014 · 2031
Closest in time.