Fetching the paper…
Reading the bibliography…
This paper presents a comprehensive survey of Federated Reinforcement Learning (FRL), an emerging and promising field in Reinforcement Learning (RL).
1901
Earlier work this paper cites.
1901
Earlier work this paper cites.
1902
Earlier work this paper cites.
1906
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1908
Earlier work this paper cites.
1910
Earlier work this paper cites.
1912
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of The 33rd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. F. Balcan and K. Q. Weinberger, Eds., vol. 48. New York, New York, USA: PMLR, 20–22 Jun 2016, pp. 1928–1937. [Online]. Available: https://proceedings.mlr.press/v48/mniha16.html
1937
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of The 33rd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. F. Balcan and K. Q. Weinberger, Eds., vol. 48. New York, New York, USA: PMLR, 20–22 Jun 2016, pp. 1928–1937. [Online]. Available: https://proceedings.mlr.press/v48/mniha16.html
1937
Earlier work this paper cites.
G. E. Monahan, “State of the art—a survey of partially observable markov decision processes: theory, models, and algorithms,” Management science , vol. 28, no. 1, pp. 1–16, 1982
1982
Earlier work this paper cites.
H. Samet, “The quadtree and related hierarchical data structures,” ACM Comput. Surv. , vol. 16, no. 2, p. 187–260, Jun. 1984. [Online]. Available: https://doi.org/10.1145/356924.356930
1984
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , vol. 8, no. 3-4, pp. 279–292, 1992. [Online]. Available: https://link.springer.com/content/pdf/10.1007/BF00992698.pdf
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
M. Tan, “Multi-agent reinforcement learning: Independent vs. cooperative agents,” in Proceedings of the tenth international conference on machine learning , 1993, pp. 330–337
1993
Earlier work this paper cites.
J. Peng and R. J. Williams, “Incremental multi-step q-learning,” in Machine Learning Proceedings 1994 . Elsevier, 1994, pp. 226–232
1994
Earlier work this paper cites.
T. L. Thorpe, “Vehicle traffic light control using sarsa,” in Online]. Available: citeseer. ist. psu. edu/thorpe97vehicle. html . Citeseer, 1997. [Online]. Available: https://citeseer.ist.psu.edu/thorpe97vehicle.html
1997
Earlier work this paper cites.
C. Szepesvári and M. L. Littman, “A unified analysis of value-function-based reinforcement-learning algorithms,” Neural computation , vol. 11, no. 8, pp. 2017–2060, 1999
1999
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in neural information processing systems , 2000, pp. 1008–1014. [Online]. Available: https://proceedings.neurips.cc/paper/1786-actor-critic-algorithms.pdf
2000
Earlier work this paper cites.
P. Stone and M. Veloso, “Multiagent systems: A survey from a machine learning perspective,” Autonomous Robots , vol. 8, no. 3, pp. 345–383, 2000
2000
Earlier work this paper cites.
M. Lauer and M. Riedmiller, “An algorithm for distributed reinforcement learning in cooperative multi-agent systems,” in In Proceedings of the Seventeenth International Conference on Machine Learning . Citeseer, 2000. [Online]. Available: http://citeseerx.ist.psu.edu/viewdoc/summary
2000
Earlier work this paper cites.
M. L. Littman, “Value-function reinforcement learning in markov games,” Cognitive systems research , vol. 2, no. 1, pp. 55–66, 2001
2001
Earlier work this paper cites.
D. S. Bernstein, R. Givan, N. Immerman, and S. Zilberstein, “The complexity of decentralized control of markov decision processes,” Mathematics of operations research , vol. 27, no. 4, pp. 819–840, 2002
2002
Earlier work this paper cites.
2003
Earlier work this paper cites.
M. Grounds and D. Kudenko, “Parallel reinforcement learning with linear function approximation,” in Adaptive Agents and Multi-Agent Systems III. Adaptation and Multi-Agent Learning , K. Tuyls, A. Nowe, Z. Guessoum, and D. Kudenko, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2008, pp. 60–74
2008
Earlier work this paper cites.
L. Busoniu, R. Babuska, and B. De Schutter, “A comprehensive survey of multiagent reinforcement learning,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) , vol. 38, no. 2, pp. 156–172, 2008
2008
Earlier work this paper cites.
2009
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on Knowledge and Data Engineering , vol. 22, no. 10, pp. 1345–1359, 2010
2010
Earlier work this paper cites.
L. Buşoniu, R. Babuška, and B. De Schutter, “Multi-agent reinforcement learning: An overview,” Innovations in multi-agent systems and applications-1 , pp. 183–221, 2010
2010
Earlier work this paper cites.
M. E. Taylor, “Teaching reinforcement learning with mario: An argument and case study,” in Second AAAI Symposium on Educational Advances in Artificial Intelligence , 2011. [Online]. Available: https://www.aaai.org/ocs/index.php/EAAI/EAAI11/paper/viewPaper/3515
2011
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, E. P. Xing and T. Jebara, Eds., vol. 32, no. 1. Bejing, China: PMLR, 22–24 Jun 2014, pp. 387–395. [Online]. Available: https://proceedings.mlr.press/v32/silver14.html
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, F. Bach and D. Blei, Eds., vol. 37. Lille, France: PMLR, 07–09 Jul 2015, pp. 1889–1897. [Online]. Available: https://proceedings.mlr.press/v37/schulman15.html
2015
Earlier work this paper cites.
M. Hausknecht and P. Stone, “Deep recurrent q-learning for partially observable mdps,” in 2015 aaai fall symposium series , 2015. [Online]. Available: https://www.aaai.org/ocs/index.php/FSS/FSS15/paper/viewPaper/11673
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015. [Online]. Available: https://doi.org/10.1038/nature14236
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 30, no. 1, 2016. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/10295
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
E. Van der Pol and F. A. Oliehoek, “Coordinated deep reinforcement learners for traffic light control,” Proceedings of Learning, Inference and Control of Multi-Agent Systems (at NIPS 2016) , 2016. [Online]. Available: https://www.elisevanderpol.nl/papers/vanderpolNIPSMALIC2016.pdf
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. Sukhbaatar, a. szlam, and R. Fergus, “Learning multiagent communication with backpropagation,” in Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available: https://proceedings.neurips.cc/paper/2016/file/55b1927fdafef39c48e5b73b5d61ea60-Paper.pdf
T. Nishio and R. Yonetani, “Client selection for federated learning with heterogeneous resources in mobile edge,” in ICC 2019 - 2019 IEEE International Conference on Communications (ICC) , 2019, pp. 1–7
2019
Later among the works it cites.
W. Y. B. Lim, N. C. Luong, D. T. Hoang, Y. Jiao, Y.-C. Liang, Q. Yang, D. Niyato, and C. Miao, “Federated learning in mobile edge networks: A comprehensive survey,” IEEE Communications Surveys Tutorials , vol. 22, no. 3, pp. 2031–2063, 2020
2020
Later among the works it cites.
T. Li, A. K. Sahu, A. Talwalkar, and V. Smith, “Federated learning: Challenges, methods, and future directions,” IEEE Signal Processing Magazine , vol. 37, no. 3, pp. 50–60, 2020
2020
Later among the works it cites.
H. Zhu and Y. Jin, “Multi-objective evolutionary federated learning,” IEEE Transactions on Neural Networks and Learning Systems , vol. 31, no. 4, pp. 1310–1322, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. E. Sallab, M. Abdou, E. Perot, and S. Yogamani, “Deep reinforcement learning framework for autonomous driving,” Electronic Imaging , vol. 2017, no. 19, pp. 70–76, 2017. [Online]. Available: 10.2352/ISSN.2470-1173.2017.19.AVM-023
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
J. Foerster, N. Nardelli, G. Farquhar, T. Afouras, P. H. S. Torr, P. Kohli, and S. Whiteson, “Stabilising experience replay for deep multi-agent reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017, pp. 1146–1155. [Online]. Available: https://proceedings.mlr.press/v70/foerster17b.html
2017
Cited alongside, same era.
2017
Cited alongside, same era.
X. Xiong, K. Zheng, L. Lei, and L. Hou, “Resource allocation based on deep reinforcement learning in iot edge computing,” IEEE Journal on Selected Areas in Communications , vol. 38, no. 6, pp. 1133–1146, 2020
2020
Later among the works it cites.
L. Lei, Y. Tan, K. Zheng, S. Liu, K. Zhang, and X. Shen, “Deep reinforcement learning for autonomous internet of things: Model, applications and challenges,” IEEE Communications Surveys Tutorials , vol. 22, no. 3, pp. 1722–1760, 2020
2020
Later among the works it cites.
X. Wang, C. Wang, X. Li, V. C. M. Leung, and T. Taleb, “Federated deep reinforcement learning for internet of things with decentralized cooperative edge caching,” IEEE Internet of Things Journal , vol. 7, no. 10, pp. 9441–9455, 2020
2020
Later among the works it cites.
W. Mao, K. Zhang, E. Miehling, and T. Başar, “Information state embedding in partially observable cooperative multi-agent reinforcement learning,” in 2020 59th IEEE Conference on Decision and Control (CDC) , 2020, pp. 6124–6131
2020
Later among the works it cites.
Y.-J. Liu, G. Feng, Y. Sun, S. Qin, and Y.-C. Liang, “Device association for ran slicing based on hybrid federated deep reinforcement learning,” IEEE Transactions on Vehicular Technology , vol. 69, no. 12, pp. 15 731–15 745, 2020
2020
Later among the works it cites.
L. Zhang, H. Yin, Z. Zhou, S. Roy, and Y. Sun, “Enhancing wifi multiple access performance with federated deep reinforcement learning,” in 2020 IEEE 92nd Vehicular Technology Conference (VTC2020-Fall) , 2020, pp. 1–6
2020
Later among the works it cites.
X. Zhang, M. Peng, S. Yan, and Y. Sun, “Deep-reinforcement-learning-based mode selection and resource allocation for cellular v2x communications,” IEEE Internet of Things Journal , vol. 7, no. 7, pp. 6380–6391, 2020
2020
Later among the works it cites.
D. Kwon, J. Jeon, S. Park, J. Kim, and S. Cho, “Multiagent ddpg-based deep learning for smart ocean federated learning iot networks,” IEEE Internet of Things Journal , vol. 7, no. 10, pp. 9895–9903, 2020
2020
Later among the works it cites.
H.-K. Lim, J.-B. Kim, J.-S. Heo, and Y.-H. Han, “Federated Reinforcement Learning for Training Control Policies on Multiple IoT Devices,” Sensors , vol. 20, no. 5, p. 1359, Mar. 2020. [Online]. Available: https://www.mdpi.com/1424-8220/20/5/1359
2020
Later among the works it cites.
N. I. Mowla, N. H. Tran, I. Doh, and K. Chae, “Afrl: Adaptive federated reinforcement learning for intelligent jamming defense in fanet,” Journal of Communications and Networks , vol. 22, no. 3, pp. 244–258, 2020
2020
Later among the works it cites.
S. Lee and D.-H. Choi, “Federated reinforcement learning for energy management of multiple smart homes with distributed energy resources,” IEEE Transactions on Industrial Informatics , pp. 1–1, 2020
2020
Later among the works it cites.
M. K. Abdel-Aziz, S. Samarakoon, C. Perfecto, and M. Bennis, “Cooperative perception in vehicular networks using multi-agent reinforcement learning,” in 2020 54th Asilomar Conference on Signals, Systems, and Computers , 2020, pp. 408–412
2020
Later among the works it cites.
H. Wang, Z. Kaplan, D. Niu, and B. Li, “Optimizing Federated Learning on Non-IID Data with Reinforcement Learning,” in IEEE INFOCOM 2020 - IEEE Conference on Computer Communications . Toronto, ON, Canada: IEEE, Jul. 2020, pp. 1698–1707. [Online]. Available: https://ieeexplore.ieee.org/document/9155494/
2020
Later among the works it cites.
H. Yu, Z. Liu, Y. Liu, T. Chen, M. Cong, X. Weng, D. Niyato, and Q. Yang, “A fairness-aware incentive scheme for federated learning,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , ser. AIES ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 393–399. [Online]. Available: https://doi.org/10.1145/3375627.3375840
2020
Later among the works it cites.
D. C. Nguyen, M. Ding, P. N. Pathirana, A. Seneviratne, J. Li, and H. Vincent Poor, “Federated learning for internet of things: A comprehensive survey,” IEEE Communications Surveys Tutorials , vol. 23, no. 3, pp. 1622–1658, 2021
2021
Closest in time.
L. U. Khan, W. Saad, Z. Han, E. Hossain, and C. S. Hong, “Federated learning for internet of things: Recent advances, taxonomy, and open challenges,” IEEE Communications Surveys Tutorials , vol. 23, no. 3, pp. 1759–1799, 2021
2021
Closest in time.
L. Lei, Y. Tan, G. Dahlenburg, W. Xiang, and K. Zheng, “Dynamic energy dispatch based on deep reinforcement learning in iot-driven smart isolated microgrids,” IEEE Internet of Things Journal , vol. 8, no. 10, pp. 7938–7953, 2021
2021
Closest in time.
L. Canese, G. C. Cardarilli, L. Di Nunzio, R. Fazzolari, D. Giardino, M. Re, and S. Spanò, “Multi-agent reinforcement learning: A review of challenges and applications,” Applied Sciences , vol. 11, no. 11, p. 4948, 2021. [Online]. Available: https://doi.org/10.3390/app11114948
2021
Closest in time.
K. Zhang, Z. Yang, and T. Başar, “Multi-agent reinforcement learning: A selective overview of theories and algorithms,” Handbook of Reinforcement Learning and Control , pp. 321–384, 2021
2021
Closest in time.
H.-R. Lee and T. Lee, “Multi-agent reinforcement learning algorithm to solve a partially-observable multi-agent problem in disaster response,” European Journal of Operational Research , vol. 291, no. 1, pp. 296–308, 2021
2021
Closest in time.
Y. Hu, Y. Hua, W. Liu, and J. Zhu, “Reward shaping based federated reinforcement learning,” IEEE Access , vol. 9, pp. 67 259–67 267, 2021
2021
Closest in time.
2021
Closest in time.
X. Wang, R. Li, C. Wang, X. Li, T. Taleb, and V. C. M. Leung, “Attention-weighted federated deep reinforcement learning for device-to-device assisted heterogeneous collaborative edge caching,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 1, pp. 154–169, 2021
2021
Closest in time.
M. Zhang, Y. Jiang, F.-C. Zheng, M. Bennis, and X. You, “Cooperative edge caching via federated deep reinforcement learning in fog-rans,” in 2021 IEEE International Conference on Communications Workshops (ICC Workshops) , 2021, pp. 1–6
2021
Closest in time.
F. Majidi, M. R. Khayyambashi, and B. Barekatain, “Hfdrl: An intelligent dynamic cooperate cashing method based on hierarchical federated deep reinforcement learning in edge-enabled iot,” IEEE Internet of Things Journal , pp. 1–1, 2021
2021
Closest in time.
L. Zhao, Y. Ran, H. Wang, J. Wang, and J. Luo, “Towards cooperative caching for vehicular networks with multi-level federated reinforcement learning,” in ICC 2021 - IEEE International Conference on Communications , 2021, pp. 1–6
2021
Closest in time.
Z. Zhu, S. Wan, P. Fan, and K. B. Letaief, “Federated multi-agent actor-critic learning for age sensitive mobile edge computing,” IEEE Internet of Things Journal , pp. 1–1, 2021
2021
Closest in time.
Z. Tianqing, W. Zhou, D. Ye, Z. Cheng, and J. Li, “Resource allocation in iot edge computing via concurrent federated reinforcement learning,” IEEE Internet of Things Journal , pp. 1–1, 2021
2021
Closest in time.
H. Huang, C. Zeng, Y. Zhao, G. Min, Y. Zhu, W. Miao, and J. Hu, “Scalable orchestration of service function chains in nfv-enabled networks: A federated reinforcement learning approach,” IEEE Journal on Selected Areas in Communications , vol. 39, no. 8, pp. 2558–2571, 2021
2021
Closest in time.
Y. Cao, S.-Y. Lien, Y.-C. Liang, and K.-C. Chen, “Federated deep reinforcement learning for user access control in open radio access networks,” in ICC 2021 - IEEE International Conference on Communications , 2021, pp. 1–6
2021
Closest in time.
M. Xu, J. Peng, B. B. Gupta, J. Kang, Z. Xiong, Z. Li, and A. A. A. El-Latif, “Multi-agent federated reinforcement learning for secure incentive mechanism in intelligent cyber-physical systems,” IEEE Internet of Things Journal , pp. 1–1, 2021
2021
Closest in time.
H.-K. Lim, J.-B. Kim, I. Ullah, J.-S. Heo, and Y.-H. Han, “Federated reinforcement learning acceleration method for precise control of multiple devices,” IEEE Access , vol. 9, pp. 76 296–76 306, 2021
2021
Closest in time.
T. G. Nguyen, T. V. Phan, D. T. Hoang, T. N. Nguyen, and C. So-In, “Federated deep reinforcement learning for traffic monitoring in sdn-based iot networks,” IEEE Transactions on Cognitive Communications and Networking , pp. 1–1, 2021
2021
Closest in time.
X. Wang, S. Garg, H. Lin, J. Hu, G. Kaddoum, M. J. Piran, and M. S. Hossain, “Towards accurate anomaly detection in industrial internet-of-things using hierarchical federated learning,” IEEE Internet of Things Journal , pp. 1–1, 2021
2021
Closest in time.
P. Zhang, P. Gan, G. S. Aujla, and R. S. Batth, “Reinforcement learning for edge device selection using social attribute perception in industry 4.0,” IEEE Internet of Things Journal , pp. 1–1, 2021
2021
Closest in time.
Y. Zhan, P. Li, W. Leijie, and S. Guo, “L4l: Experience-driven computational resource control in federated learning,” IEEE Transactions on Computers , pp. 1–1, 2021
2021
Closest in time.
Y. Dong, P. Gan, G. S. Aujla, and P. Zhang, “Ra-rl: Reputation-aware edge device selection method based on reinforcement learning,” in 2021 IEEE 22nd International Symposium on a World of Wireless, Mobile and Multimedia Networks (WoWMoM) , 2021, pp. 348–353
2021
Closest in time.
M. Chen, H. V. Poor, W. Saad, and S. Cui, “Convergence time optimization for federated learning over wireless networks,” IEEE Transactions on Wireless Communications , vol. 20, no. 4, pp. 2457–2471, 2021
2021
Closest in time.
A. Anwar and A. Raychowdhury, “Multi-task federated reinforcement learning with adversaries,” 2021
2021
Closest in time.