Fetching the paper…
Reading the bibliography…
We consider the problem of federated Q-learning, where $M$ agents aim to collaboratively learn the optimal Q-function of an unknown infinite-horizon Markov decision process with finite state and action spaces.
M. J. Wainwright · 1905
Earlier work this paper cites.
Variance-reduced Q-learning is minimax optimal
M. J. Wainwright · 1906
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 1937
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 1937
Earlier work this paper cites.
On tail probabilities for martingales
D. A. Freedman · 1975
Earlier work this paper cites.
Communication complexity of convex optimization
J. N. Tsitsiklis and Z. Q. Luo · 1987
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
T. Jaakkola, M. Jordan, and S. Singh · 1993
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
The asymptotic convergence-rate of Q-learning
C. Szepesvári · 1997
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
M. Kearns and S. Singh · 1998
Earlier work this paper cites.
The O.D.E. method for convergence of stochastic approximation and reinforcement learning
V. S. Borkar and S. P. Meyn · 2000
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2001
Earlier work this paper cites.
Learning rates for Q-learning
E. Even-Dar and Y. Mansour · 2004
Earlier work this paper cites.
Double Q-learning
H. v. Hasselt · 2010
Earlier work this paper cites.
Error bounds for constant step-size Q-learning
C. Beck and R. Srikant · 2012
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
M. G. Azar, R. Munos, and H. J. Kappen · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
Optimality guarantees for distributed statistical estimation
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and Y. Zhang · 2014
Earlier work this paper cites.
Markov decision processes: Discrete Stochastic Dynamic Programming
M. Puterman · 2014
Earlier work this paper cites.
Communication lower bounds for statistical estimation problems via a distributed data processing inequality
M. Braverman, A. Garg, T. Ma, H. L. Nguyen, and D. P. Woodruff · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershalvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Cited alongside, same era.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, and K. Kavukcuoglu · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
A. Sidford, M. Wang, X. Wu, L. Yang, and Y. Ye · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. Sutton and A. Barton · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
Federated Multi-Armed Bandits
C. Shi and C. Shen · 2021
Later among the works it cites.
The min-max complexity of distributed stochastic convex optimization with intermittent communication
B. Woodworth, B. Bullins, O. Shamir, and N. Srebro · 2021
Later among the works it cites.
Sample and communication-efficient decentralized actor-critic algorithms with finite-time analysis
Z. Chen, Y. Zhou, R.-R. Chen, and S. Zou · 2022
Later among the works it cites.
Federated reinforcement learning with environment heterogeneity
H. Jin, Y. Peng, W. Yang, S. Wang, and Z. Zhang · 2022
Later among the works it cites.
Federated reinforcement learning: Linear speedup under markovian sampling
S. Khodadadian, P. Sharma, G. Joshi, and S. T. Maguluri · 2022
Later among the works it cites.
Pessimistic Q-learning for offline reinforcement learning: Towards optimal sample complexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Vershynin · 2018
Cited alongside, same era.
Graph oracle models, lower bounds, and gaps for parallel stochastic optimization
B. Woodworth, J. Wang, A. Smith, B. McMahan, and N. Srebro · 2018
Cited alongside, same era.
Gossip-based actor-learner architectures for deep reinforcement learning
M. Assran, J. Romoff, N. Ballas, J. Pineau, and M. Rabbat · 2019
Cited alongside, same era.
Provably efficient Q-learning with low switching cost
Y. Bai, T. Xie, N. Jiang, and Y.-X. Wang · 2019
Cited alongside, same era.
Finite-time analysis of distributed TD(0) with linear function approximation on multi-agent reinforcement learning
T. Doan, S. Maguluri, and J. Romberg · 2019
Cited alongside, same era.
Local SGD with periodic averaging: Tighter analysis and adaptive synchronization
F. Haddadpour, M. M. Kamani, M. Mahdavi, and V. R. Cadambe · 2019
Cited alongside, same era.
Finite-sample analysis of contractive stochastic approximation using smooth convex envelopes
Z. Chen, S. T. Maguluri, S. Shakkottai, and K. Shanmugam · 2020
Cited alongside, same era.
L. Shi, G. Li, Y. Wei, Y. Chen, and Y. Chi · 2022
Later among the works it cites.
Distributed TD(0) with almost no communication
R. Liu and A. Olshevsky · 2023
Later among the works it cites.
Distributed linear bandits under communication constraints
S. Salgia and Q. Zhao · 2023
Later among the works it cites.
Towards understanding asynchronous advantage actor-critic: Convergence and linear speedup
H. Shen, K. Zhang, M. Hong, and T. Chen · 2023
Later among the works it cites.
H. Wang, A. Mitra, H. Hassani, G. J. Pappas, and J. Anderson · 2023
Later among the works it cites.
The blessing of heterogeneity in federated q-learning: Linear speedup and beyond
J. Woo, G. Joshi, and Y. Chi · 2023
Later among the works it cites.
Fedkl: Tackling data heterogeneity in federated reinforcement learning by penalizing kl divergence
Z. Xie and S. Song · 2023
Later among the works it cites.
The efficacy of pessimism in asynchronous Q-learning
Y. Yan, G. Li, Y. Chen, and J. Fan · 2023
Later among the works it cites.
Federated natural policy gradient and actor critic methods for multi-task reinforcement learning
T. Yang, S. Cen, Y. Wei, Y. Chen, and Y. Chi · 2023
Later among the works it cites.
G. Lan, D.-J. Han, A. Hashemi, V. Aggarwal, and C. G. Brinton · 2024
Closest in time.
Is Q-learning minimax optimal? a tight sample complexity analysis
G. Li, C. Cai, Y. Chen, Y. Wei, and Y. Chi · 2024
Closest in time.
One-shot averaging for distributed TD ( λ \lambda ) under Markov sampling
H. Tian, I. C. Paschalidis, and A. Olshevsky · 2024
Closest in time.
Federated offline reinforcement learning: Collaborative single-policy coverage suffices
J. Woo, L. Shi, G. Joshi, and Y. Chi · 2024
Closest in time.
Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning
E. Xia, K. Khamaru, M. J. Wainwright, and M. I. Jordan · 2024
Closest in time.
Finite-time analysis of on-policy heterogeneous federated reinforcement learning, 2024
C. Zhang, H. Wang, A. Mitra, and J. Anderson · 2024
Closest in time.
Federated Q-learning: Linear regret speedup with low communication cost
Z. Zheng, F. Gao, L. Xue, and J. Yang · 2024
Closest in time.