Fetching the paper…
Reading the bibliography…
Federated reinforcement learning (RL) enables collaborative decision making of multiple distributed agents without sharing local data trajectories.
Federated deep reinforcement learning
Zhuo, H. H., Feng, W., Lin, Y., Xu, Q., and Yang, Q. (2019) · 1901
Earlier work this paper cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z. (2019) · 1909
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Quantal response equilibria for normal form games
McKelvey, R. D. and Palfrey, T. R. (1995) · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2001) · 2001
Earlier work this paper cites.
Improving sample complexity bounds for actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y. (2020) · 2004
Earlier work this paper cites.
Convergence analysis of distributed subgradient methods over random networks
Lobel, I. and Ozdaglar, A. (2008) · 2008
Earlier work this paper cites.
The matrix cookbook
Petersen, K. B. and Pedersen, M. S. (2008) · 2008
Earlier work this paper cites.
Natural actor-critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Nedic, A. and Ozdaglar, A. (2009) · 2009
Earlier work this paper cites.
Discrete-time dynamic average consensus
Zhu, M. and Martínez, S. (2010) · 2010
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
Duchi, J. C., Agarwal, A., and Wainwright, M. J. (2011) · 2011
Earlier work this paper cites.
Matrix analysis
Horn, R. A. and Johnson, C. R. (2012) · 2012
Earlier work this paper cites.
Kar, S., Moura, J. M., and Poor, H. V. (2012) · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Bach, F. and Moulines, E. (2013) · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Earlier work this paper cites.
Next: In-network nonconvex optimization
Di Lorenzo, P. and Scutari, G. (2016) · 2016
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J. (2017) · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Cited alongside, same era.
Achieving geometric convergence for distributed optimization over time-varying graphs
Nedic, A., Olshevsky, A., and Shi, W. (2017) · 2017
Cited alongside, same era.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J. (2017) · 2017
Cited alongside, same era.
Harnessing smoothness to accelerate distributed optimization
Qu, G. and Li, N. (2017) · 2017
Cited alongside, same era.
Fast policy extragradient methods for competitive games with entropy regularization
Cen, S., Wei, Y., and Chi, Y. (2021) · 2021
Later among the works it cites.
Multi-agent off-policy TDC with near-optimal sample and communication complexity
Chen, Z., Zhou, Y., and Chen, R. (2021b) · 2021
Later among the works it cites.
Maximum entropy RL (provably) solves some robust RL problems
Eysenbach, B. and Levine, S. (2021) · 2021
Later among the works it cites.
On the linear convergence of natural policy gradient algorithm
Khodadadian, S., Jhunjhunwala, P. R., Varma, S. M., and Maguluri, S. T. (2021) · 2021
Later among the works it cites.
Distributed stochastic gradient tracking methods
Pu, S. and Nedić, A. (2021) · 2021
Later among the works it cites.
Federated reinforcement learning: Techniques, applications, and open challenges
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. (2018) · 2018
Cited alongside, same era.
Network topology and communication-computation tradeoffs in decentralized optimization
Nedić, A., Olshevsky, A., and Rabbat, M. G. (2018) · 2018
Cited alongside, same era.
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D. (2019) · 2019
Cited alongside, same era.
Gossip-based actor-learner architectures for deep reinforcement learning
Assran, M., Romoff, J., Ballas, N., Pineau, J., and Rabbat, M. (2019) · 2019
Cited alongside, same era.
Communication-efficient distributed optimization in networks with gradient tracking and variance reduction
Li, B., Cen, S., Chen, Y., and Chi, Y. (2020) · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D. (2020) · 2020
Cited alongside, same era.
Qi, J., Zhou, Q., Lei, L., and Zheng, K. (2021) · 2021
Later among the works it cites.
A decentralized policy gradient approach to multi-task reinforcement learning
Zeng, S., Anwar, M. A., Doan, T. T., Raychowdhury, A., and Romberg, J. (2021) · 2021
Later among the works it cites.
Chen, J., Feng, J., Gao, W., and Wei, K. (2022) · 2022
Later among the works it cites.
Exploring the role of artificial intelligence in enhancing academic performance: A case study of chatgpt
M Alshater, M. (2022) · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Xiao, L. (2022) · 2022
Later among the works it cites.
Linear convergence of natural policy gradient methods with log-linear policies
Yuan, R., Du, S. S., Gower, R. M., Lazaric, A., and Xiao, L. (2022) · 2022
Later among the works it cites.
Anchor-changing regularized natural policy gradient for multi-objective reinforcement learning
Zhou, R., Liu, T., Kalathil, D., Kumar, P., and Tian, C. (2022) · 2022
Later among the works it cites.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Lan, G. (2023) · 2023
Closest in time.
Chatgpt and academic research: a review and recommendations based on practical examples
Rahman, M. M., Terano, H. J., Rahman, M. N., Salamzadeh, A., and Rahaman, M. S. (2023) · 2023
Closest in time.
Federated ensemble model-based reinforcement learning in edge computing
Wang, J., Hu, J., Mills, J., Min, G., Xia, M., and Georgalas, N. (2023) · 2023
Closest in time.
The blessing of heterogeneity in federated q-learning: Linear speedup and beyond
Woo, J., Joshi, G., and Chi, Y. (2023) · 2023
Closest in time.
Policy mirror descent for regularized reinforcement learning: A generalized framework with linear convergence
Zhan, W., Cen, S., Huang, B., Chen, Y., Lee, J. D., and Chi, Y. (2023) · 2023
Closest in time.
Federated multi-objective reinforcement learning
Zhao, F., Ren, X., Yang, S., Zhao, P., Zhang, R., and Xu, X. (2023) · 2023
Closest in time.
Federated offline reinforcement learning: Collaborative single-policy coverage suffices
Woo, J., Shi, L., Joshi, G., and Chi, Y. (2024) · 2024
Closest in time.