Fetching the paper…
Reading the bibliography…
Traffic optimization challenges, such as load balancing, flow scheduling, and improving packet delivery time, are difficult online decision-making problems in wide area networks (WAN).
A note on dijkstra’s shortest path algorithm
Donald B Johnson · 1973
Earlier work this paper cites.
A loop-free extended bellman-ford routing protocol without bouncing effect
Chunhsiang Cheng, Ralph Riley, Srikanta PR Kumar, and Jose J Garcia-Luna-Aceves · 1989
Earlier work this paper cites.
Packet routing in dynamically changing networks: A reinforcement learning approach
Justin A Boyan and Michael L Littman · 1994
Earlier work this paper cites.
Computer networks
Andrew S Tanenbaum et al · 1996
Earlier work this paper cites.
Predictive q-routing: A memory-based reinforcement learning approach to adaptive traffic control
Samuel PM Choi and Dit-Yan Yeung · 1996
Earlier work this paper cites.
Dual reinforcement q-routing: An online adaptive routing algorithm
Shailesh Kumar and Risto Miikkulainen · 1997
Earlier work this paper cites.
Ants and reinforcement learning: A case study in routing in dynamic networks
Devika Subramanian, Peter Druschel, and Johnny Chen · 1997
Earlier work this paper cites.
Flooding in wireless ad hoc networks
Hyojun Lim and Chongkwon Kim · 2001
Earlier work this paper cites.
A multi-agent, policy-gradient approach to network routing
Nigel Tao, Jonathan Baxter, and Lex Weaver · 2001
Earlier work this paper cites.
Overview and principles of internet traffic engineering
D. Awduche, A. Chiu, A. Elwalid, I. Widjaja, and X. Xiao · 2002
Earlier work this paper cites.
Reinforcement learning for adaptive routing
Leonid Peshkin and Virginia Savova · 2002
Earlier work this paper cites.
A review of routing protocols for mobile ad hoc networks
Mehran Abolhasan, Tadeusz Wysocki, and Eryk Dutkiewicz · 2004
Earlier work this paper cites.
Evolution strategies in dynamic environments
Lutz Schönemann · 2007
Earlier work this paper cites.
Hysteretic q-learning: an algorithm for decentralized reinforcement learning in cooperative multi-agent teams
Laëtitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2007
Earlier work this paper cites.
A survey on key management mechanisms for distributed wireless sensor networks
Marcos A Simplício Jr, Paulo SLM Barreto, Cintia B Margi, and Tereza CMB Carvalho · 2010
Earlier work this paper cites.
Simulated annealing based hierarchical q-routing: a dynamic routing protocol
Antonio Mira Lopez and Douglas R Heisterkamp · 2011
Cited alongside, same era.
B4: Experience with a globally-deployed software defined wan
Sushant Jain, Alok Kumar, Subhasree Mandal, Joon Ong, Leon Poutievski, Arjun Singh, Subbaiah Venkata, Jim Wanderer, Junlan Zhou, Min Zhu, et al · 2013
Cited alongside, same era.
Achieving high utilization with software-driven wan
Chi-Yao Hong, Srikanth Kandula, Ratul Mahajan, Ming Zhang, Vijay Gill, Mohan Nanduri, and Roger Wattenhofer · 2013
Cited alongside, same era.
Application of reinforcement learning to routing in distributed wireless networks: a review
Hasan AA Al-Rawi, Ming Ann Ng, and Kok-Lim Alvin Yau · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch · 2017
Later among the works it cites.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip HS Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Later among the works it cites.
A comprehensive survey on machine learning for networking: evolution, applications and research opportunities
Raouf Boutaba, Mohammad A Salahuddin, Noura Limam, Sara Ayoubi, Nashid Shahriar, Felipe Estrada-Solano, and Oscar M Caicedo · 2018
Later among the works it cites.
B4 and after: managing hierarchy, partitioning, and asymmetry for availability and scale in google’s software-defined wan
Chi-Yao Hong, Subhasree Mandal, Mohammad Al-Fares, Min Zhu, Richard Alimi, Chandan Bhagat, Sourabh Jain, Jay Kaimal, Shiyu Liang, Kirill Mendelev, et al · 2018
Later among the works it cites.
Experience-driven networking: A deep reinforcement learning based approach
Zhiyuan Xu, Jian Tang, Jingsong Meng, Weiyi Zhang, Yanzhi Wang, Chi Harold Liu, and Dejun Yang · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Resource management with deep reinforcement learning
Hongzi Mao, Mohammad Alizadeh, Ishai Menache, and Srikanth Kandula · 2016
Cited alongside, same era.
Evolve or die: High-availability design principles drawn from googles network infrastructure
Ramesh Govindan, Ina Minei, Mahesh Kallahalla, Bikash Koley, and Amin Vahdat · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Cited alongside, same era.
Huaguang Zhang, He Jiang, Yanhong Luo, and Geyang Xiao · 2016
Cited alongside, same era.
Learning to route
Asaf Valadarsky, Michael Schapira, Dafna Shahaf, and Aviv Tamar · 2017
Cited alongside, same era.
Later among the works it cites.
Networked multi-agent reinforcement learning in continuous spaces
Kaiqing Zhang, Zhuoran Yang, and Tamer Basar · 2018
Later among the works it cites.
Multi-agent reinforcement learning via double averaging primal-dual optimization
Hoi-To Wai, Zhuoran Yang, Zhaoran Wang, and Mingyi Hong · 2018
Later among the works it cites.
Teavar: striking the right utilization-availability balance in wan traffic engineering
Jeremy Bogle, Nikhil Bhatia, Manya Ghobadi, Ishai Menache, Nikolaj Bjørner, Asaf Valadarsky, and Michael Schapira · 2019
Later among the works it cites.
Tutorial on dynamic average consensus: The problem, its applications, and the algorithms
Solmaz S Kia, Bryan Van Scoy, Jorge Cortes, Randy A Freeman, Kevin M Lynch, and Sonia Martinez · 2019
Later among the works it cites.
Multi-agent deep learning for simultaneous optimization for time and energy in distributed routing system
Dmitry Mukhutdinov, Andrey Filchenkov, Anatoly Shalyto, and Valeriy Vyatkin · 2019
Later among the works it cites.
Actor-attention-critic for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Later among the works it cites.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Shihui Li, Yi Wu, Xinyue Cui, Honghua Dong, Fei Fang, and Stuart Russell · 2019
Later among the works it cites.
Routing on multiple optimality criteria
João Luís Sobrinho and Miguel Alves Ferreira · 2020
Later among the works it cites.
Application of deep reinforcement learning in traffic signal control: An overview and impact of open traffic data
Martin Gregurić, Miroslav Vujić, Charalampos Alexopoulos, and Mladen Miletić · 2020
Later among the works it cites.
Toward packet routing with fully distributed multiagent deep reinforcement learning
Xinyu You, Xuanjie Li, Yuedong Xu, Hui Feng, Jin Zhao, and Huaicheng Yan · 2020
Later among the works it cites.