Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (DRL) has empowered a variety of artificial intelligence fields, including pattern recognition, robotics, recommendation-systems, and gaming.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning . PMLR, 2016, pp. 1928–1937
1937
Earlier work this paper cites.
S. Muggleton and L. De Raedt, “Inductive logic programming: Theory and methods,” The Journal of Logic Programming , vol. 19, pp. 629–679, 1994
1994
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” Advances in neural information processing systems , vol. 12, 1999
1999
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International conference on machine learning . PMLR, 2016, pp. 1995–2003
2003
Earlier work this paper cites.
T. Yamada, “Studies on metaheuristics for jobshop and flowshop scheduling problems,” 2003
2003
Earlier work this paper cites.
D. Srinivasan, M. C. Choy, and R. L. Cheu, “Neural networks for real-time traffic signal control,” IEEE Transactions on intelligent transportation systems , vol. 7, no. 3, pp. 261–272, 2006
2006
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE transactions on neural networks , vol. 20, no. 1, pp. 61–80, 2008
2008
Earlier work this paper cites.
S. Sanner et al. , “Relational dynamic influence diagram language (rddl): Language description,” Unpublished ms. Australian National University , vol. 32, p. 27, 2010
2010
Earlier work this paper cites.
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska, “A survey of actor-critic reinforcement learning: Standard and natural policy gradients,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) , vol. 42, no. 6, pp. 1291–1307, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning . PMLR, 2015, pp. 1889–1897
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International conference on machine learning . PMLR, 2016, pp. 49–58
2016
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 30, no. 1, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. Dai, B. Dai, and L. Song, “Discriminative embeddings of latent variable models for structured data,” in International conference on machine learning . PMLR, 2016, pp. 2702–2711
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
“Mit technology review,” https://www.technologyreview.com/10-breakthrough-technologies/ , 2017
2017
Earlier work this paper cites.
Q. Wang, Z. Mao, B. Wang, and L. Guo, “Knowledge graph embedding: A survey of approaches and applications,” IEEE Transactions on Knowledge and Data Engineering , vol. 29, no. 12, pp. 2724–2743, 2017
2017
Earlier work this paper cites.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274 , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
W. Hamilton, Z. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Khalil, H. Dai, Y. Zhang, B. Dilkina, and L. Song, “Learning combinatorial optimization algorithms over graphs,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
S. Hou, Y. Ye, Y. Song, and M. Abdulhayoglu, “Hindroid: An intelligent android malware detection system based on structured heterogeneous information network,” in Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining , 2017, pp. 1507–1515
2017
Earlier work this paper cites.
A. Bernstein and E. Burnaev, “Reinforcement learning in computer vision,” in Tenth International Conference on Machine Vision (ICMV 2017) , vol. 10696. SPIE, 2018, pp. 458–464
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
D. Zügner, A. Akbarnejad, and S. Günnemann, “Adversarial attacks on neural networks for graph data,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2018, pp. 2847–2856
2018
Earlier work this paper cites.
H. Dai, H. Li, T. Tian, X. Huang, L. Wang, J. Zhu, and L. Song, “Adversarial attack on graph structured data,” in International conference on machine learning . PMLR, 2018, pp. 1115–1124
2018
Earlier work this paper cites.
J. Jin and X. Ma, “Hierarchical multi-agent control of traffic lights based on collective learning,” Engineering applications of artificial intelligence , vol. 68, pp. 236–248, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Jiang and Z. Lu, “Learning attentional communication for multi-agent cooperation,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Wang, R. Liao, J. Ba, and S. Fidler, “Nervenet: Learning structured policy with graph neural networks,” in International conference on learning representations , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Deudon, P. Cournut, A. Lacoste, Y. Adulyasak, and L.-M. Rousseau, “Learning heuristics for the tsp by policy gradient,” in International conference on the integration of constraint programming, artificial intelligence, and operations research . Springer, 2018, pp. 170–181
2018
Earlier work this paper cites.
D. Baumann, J.-J. Zhu, G. Martius, and S. Trimpe, “Deep reinforcement learning for event-triggered control,” in 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 2018, pp. 943–950
2018
Earlier work this paper cites.
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson, “Counterfactual multi-agent policy gradients,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
M. Everett, Y. F. Chen, and J. P. How, “Motion planning among dynamic, decision-making agents with deep reinforcement learning,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 3052–3059
2018
Earlier work this paper cites.
Y. Zhang, H. Dai, Z. Kozareva, A. J. Smola, and L. Song, “Variational reasoning for question answering with knowledge graph,” in Thirty-second AAAI conference on artificial intelligence , 2018
2018
Earlier work this paper cites.
M. Schlichtkrull, T. N. Kipf, P. Bloem, R. v. d. Berg, I. Titov, and M. Welling, “Modeling relational data with graph convolutional networks,” in European semantic web conference . Springer, 2018, pp. 593–607
2018
Earlier work this paper cites.
M. Nazari, A. Oroojlooy, L. Snyder, and M. Takác, “Reinforcement learning for solving the vehicle routing problem,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
C. Shi, B. Hu, W. X. Zhao, and S. Y. Philip, “Heterogeneous information network embedding for recommendation,” IEEE Transactions on Knowledge and Data Engineering , vol. 31, no. 2, pp. 357–370, 2018
2018
Earlier work this paper cites.
B. Hu, C. Shi, W. X. Zhao, and P. S. Yu, “Leveraging meta-path based context for top-n recommendation with a neural co-attention model,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2018, pp. 1531–1540
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Glavic, “(deep) reinforcement learning for electric power system control and related problems: A short review and perspectives,” Annual Reviews in Control , vol. 48, pp. 22–35, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Iqbal and F. Sha, “Actor-attention-critic for multi-agent reinforcement learning,” in International Conference on Machine Learning . PMLR, 2019, pp. 2961–2970
2019
Cited alongside, same era.
H. Lu, X. Zhang, and S. Yang, “A learning-based iterative method for solving vehicle routing problems,” in International conference on learning representations , 2019
2019
Cited alongside, same era.
S. Ahn, Y. Seo, and J. Shin, “Learning what to defer for maximum independent sets,” in International Conference on Machine Learning . PMLR, 2020, pp. 134–144
2020
Later among the works it cites.
L. Hu, S. Xu, C. Li, C. Yang, C. Shi, N. Duan, X. Xie, and M. Zhou, “Graph neural news recommendation with unsupervised preference disentanglement,” in Proceedings of the 58th annual meeting of the association for computational linguistics , 2020, pp. 4255–4264
2020
Later among the works it cites.
Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and S. Y. Philip, “A comprehensive survey on graph neural networks,” IEEE transactions on neural networks and learning systems , vol. 32, no. 1, pp. 4–24, 2020
2020
Later among the works it cites.
B. Chalaki, L. E. Beaver, B. Remer, K. Jang, E. Vinitsky, A. M. Bayen, and A. A. Malikopoulos, “Zero-shot autonomous vehicle policy transfer: From simulation to real-world via adversarial learning,” in 2020 IEEE 16th International Conference on Control & Automation (ICCA) . IEEE, 2020, pp. 35–40
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Yang, B. Yang, H.-S. Wong, and Z. Kang, “Cooperative traffic signal control using multi-step return and off-policy asynchronous advantage actor-critic graph algorithm,” Knowledge-Based Systems , vol. 183, p. 104855, 2019
2019
Cited alongside, same era.
Q. Xiao, C. Li, Y. Tang, and L. Li, “Meta-reinforcement learning of machining parameters for energy-efficient process control of flexible turning operations,” IEEE Transactions on Automation Science and Engineering , vol. 18, no. 1, pp. 5–18, 2019
2019
Cited alongside, same era.
F. Gama, A. G. Marques, G. Leus, and A. Ribeiro, “Convolutional neural network architectures for signals supported on graphs,” IEEE Transactions on Signal Processing , vol. 67, no. 4, pp. 1034–1049, 2019
2019
Cited alongside, same era.
X. Wang, D. Wang, C. Xu, X. He, Y. Cao, and T.-S. Chua, “Explainable reasoning over knowledge graphs for recommendation,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, no. 01, 2019, pp. 5329–5336
2019
Cited alongside, same era.
K. Do, T. Tran, and S. Venkatesh, “Graph transformation policy network for chemical reaction prediction,” in Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2019, pp. 750–760
2019
Cited alongside, same era.
J. You, R. Ying, and J. Leskovec, “Position-aware graph neural networks,” in International Conference on Machine Learning . PMLR, 2019, pp. 7134–7143
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Zhuang, N. Dvornek, X. Li, and J. S. Duncan, “Ordinary differential equations on graph networks,” 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni, “Green ai,” Commun. ACM , vol. 63, no. 12, p. 54–63, nov 2020. [Online]. Available: https://doi.org/10.1145/3381831
2020
Later among the works it cites.
N. Mazyavkina, S. Sviridov, S. Ivanov, and E. Burnaev, “Reinforcement learning for combinatorial optimization: A survey,” Computers & Operations Research , vol. 134, p. 105400, 2021
2021
Later among the works it cites.
N. P. Farazi, B. Zou, T. Ahamed, and L. Barua, “Deep reinforcement learning in transportation research: A review,” Transportation research interdisciplinary perspectives , vol. 11, p. 100425, 2021
2021
Later among the works it cites.
S. K. Zhou, H. N. Le, K. Luu, H. V. Nguyen, and N. Ayache, “Deep reinforcement learning in medical imaging: A literature review,” Medical image analysis , vol. 73, p. 102193, 2021
2021
Later among the works it cites.
M. S. Frikha, S. M. Gammar, A. Lahmadi, and L. Andrey, “Reinforcement and deep reinforcement learning for wireless internet of things: A survey,” Computer Communications , vol. 178, pp. 98–113, 2021
2021
Later among the works it cites.
R. N. Boute, J. Gijsbrechts, W. van Jaarsveld, and N. Vanvuchelen, “Deep reinforcement learning for inventory control: A roadmap,” European Journal of Operational Research , 2021
2021
Later among the works it cites.
A. Perera and P. Kamalaruban, “Applications of reinforcement learning in energy systems,” Renewable and Sustainable Energy Reviews , vol. 137, p. 110618, 2021
2021
Later among the works it cites.
F. Obite, A. D. Usman, and E. Okafor, “An overview of deep reinforcement learning for spectrum sensing in cognitive radio networks,” Digital Signal Processing , vol. 113, p. 103014, 2021
2021
Later among the works it cites.
S. Ji, S. Pan, E. Cambria, P. Marttinen, and S. Y. Philip, “A survey on knowledge graphs: Representation, acquisition, and applications,” IEEE Transactions on Neural Networks and Learning Systems , vol. 33, no. 2, pp. 494–514, 2021
2021
Later among the works it cites.
C. Shan, Y. Shen, Y. Zhang, X. Li, and D. Li, “Reinforcement learning enhanced explainer for graph neural networks,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
S. Shen, Y. Fu, H. Su, H. Pan, P. Qiao, Y. Dou, and C. Wang, “Graphcomm: A graph neural network based method for multi-agent reinforcement learning,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 3510–3514
2021
Later among the works it cites.
X. Zhang, Y. Liu, X. Xu, Q. Huang, H. Mao, and A. Carie, “Structural relational inference actor-critic for multi-agent reinforcement learning,” Neurocomputing , vol. 459, pp. 383–394, 2021
2021
Later among the works it cites.
W. J. Yun, S. Yi, and J. Kim, “Multi-agent deep reinforcement learning using attentive graph neural architectures for real-time strategy games,” in 2021 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, 2021, pp. 2967–2972
2021
Later among the works it cites.
H. Chen, W. Qiu, H.-C. Ou, B. An, and M. Tambe, “Contingency-aware influence maximization: A reinforcement learning approach,” in Uncertainty in Artificial Intelligence . PMLR, 2021, pp. 1535–1545
2021
Later among the works it cites.
E. Meirom, H. Maron, S. Mannor, and G. Chechik, “Controlling graph dynamics with reinforcement learning and graph neural networks,” in International Conference on Machine Learning . PMLR, 2021, pp. 7565–7577
2021
Later among the works it cites.
V.-A. Darvariu, S. Hailes, and M. Musolesi, “Goal-directed graph construction using reinforcement learning,” Proceedings of the Royal Society A , vol. 477, no. 2254, p. 20210168, 2021
2021
Later among the works it cites.
——, “Bayesian graph neural network for fast identification of critical nodes in uncertain complex networks,” in 2021 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, 2021, pp. 3245–3251
2021
Later among the works it cites.
F.-X. Devailly, D. Larocque, and L. Charlin, “Ig-rl: Inductive graph reinforcement learning for massive-scale traffic signal control,” IEEE Transactions on Intelligent Transportation Systems , 2021
2021
Later among the works it cites.
S. Yang, B. Yang, Z. Kang, and L. Deng, “Ihg-ma: Inductive heterogeneous graph multi-agent reinforcement learning for multi-intersection traffic signal control,” Neural networks , vol. 139, pp. 265–277, 2021
2021
Later among the works it cites.
J. Yoon, K. Ahn, J. Park, and H. Yeo, “Transferable traffic signal control: Reinforcement learning with graph centric state representation,” Transportation Research Part C: Emerging Technologies , vol. 130, p. 103321, 2021
2021
Later among the works it cites.
J. Huang, J. Zhang, Q. Chang, and R. X. Gao, “Integrated process-system modelling and control through graph neural network and reinforcement learning,” CIRP Annals , vol. 70, no. 1, pp. 377–380, 2021
2021
Later among the works it cites.
J. Park, J. Chun, S. H. Kim, Y. Kim, and J. Park, “Learning to schedule job-shop problems: representation and policy learning using graph neural network and reinforcement learning,” International Journal of Production Research , vol. 59, no. 11, pp. 3360–3377, 2021
2021
Later among the works it cites.
Y. Wang, S. Hou, and X. Wang, “Reinforcement learning-based bird-view automated vehicle control to avoid crossing traffic,” Computer-Aided Civil and Infrastructure Engineering , vol. 36, no. 7, pp. 890–901, 2021
2021
Later among the works it cites.
S. Chen, J. Dong, P. Ha, Y. Li, and S. Labi, “Graph neural network and reinforcement learning for multi-agent cooperative control of connected autonomous vehicles,” Computer-Aided Civil and Infrastructure Engineering , vol. 36, no. 7, pp. 838–857, 2021
2021
Later among the works it cites.
P. Zheng, L. Xia, C. Li, X. Li, and B. Liu, “Towards self-x cognitive manufacturing network: An industrial knowledge graph-based multi-agent reinforcement learning approach,” Journal of Manufacturing Systems , vol. 61, pp. 16–26, 2021
2021
Later among the works it cites.
H. Fei, Y. Ren, Y. Zhang, D. Ji, and X. Liang, “Enriching contextualized language model from knowledge graph for biomedical information extraction,” Briefings in bioinformatics , vol. 22, no. 3, p. bbaa110, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Heuillet, F. Couthouis, and N. Díaz-Rodríguez, “Explainability in deep reinforcement learning,” Knowledge-Based Systems , vol. 214, p. 106685, 2021
2021
Later among the works it cites.
Y. Liu, A. Halev, and X. Liu, “Policy learning with constraints in model-free reinforcement learning: A survey,” in Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence , 2021
2021
Later among the works it cites.
Q. Cappart, T. Moisan, L.-M. Rousseau, I. Prémont-Schwarz, and A. A. Cire, “Combining reinforcement learning and constraint programming for combinatorial optimization,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 5, 2021, pp. 3677–3687
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Li, G. Peng, Q. Zhu, and T. Başar, “The confluence of networks, games, and learning a game-theoretic framework for multiagent decision making over networks,” IEEE Control Systems Magazine , vol. 42, no. 4, pp. 35–67, 2022
2022
Closest in time.
A. Alomari, N. Idris, A. Q. M. Sabri, and I. Alsmadi, “Deep reinforcement and transfer learning for abstractive text summarization: A review,” Computer Speech & Language , vol. 71, p. 101276, 2022
2022
Closest in time.
Y. Zhou, H. Zheng, X. Huang, S. Hao, D. Li, and J. Zhao, “Graph neural networks: Taxonomy, advances, and trends,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 13, no. 1, pp. 1–54, 2022
2022
Closest in time.
2022
Closest in time.
S. Munikoti, L. Das, and B. Natarajan, “Scalable graph neural network-based framework for identifying critical nodes and links in complex networks,” Neurocomputing , vol. 468, pp. 211–221, 2022
2022
Closest in time.
M. Noaeen, A. Naik, L. Goodman, J. Crebo, T. Abrar, Z. S. H. Abad, A. L. Bazzan, and B. Far, “Reinforcement learning in urban network traffic signal control: A systematic literature review,” Expert Systems with Applications , p. 116830, 2022
2022
Closest in time.
P. Shang, X. Liu, C. Yu, G. Yan, Q. Xiang, and X. Mi, “A new ensemble deep graph reinforcement learning network for spatio-temporal traffic volume forecasting in a freeway network,” Digital Signal Processing , vol. 123, p. 103419, 2022
2022
Closest in time.
P. Almasan, J. Suárez-Varela, K. Rusek, P. Barlet-Ros, and A. Cabellos-Aparicio, “Deep reinforcement learning meets graph neural networks: Exploring a routing optimization use case,” Computer Communications , 2022
2022
Closest in time.
A. D. McNaughton, C. Knutson, M. Bontha, J. A. Pope, and N. Kumar, “De novo design of protein target specific scaffold-based inhibitors via reinforcement learning,” in ICLR2022 Machine Learning for Drug Discovery , 2022
2022
Closest in time.
2022
Closest in time.
L. Furieri, C. L. Galimberti, M. Zakwan, and G. Ferrari-Trecate, “Distributed neural network control with dependability guarantees: a compositional port-hamiltonian approach,” in Learning for Dynamics and Control Conference . PMLR, 2022, pp. 571–583
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
X. Wang, H. Ji, C. Shi, B. Wang, Y. Ye, P. Cui, and P. S. Yu, “Heterogeneous graph attention network,” in The world wide web conference , 2019, pp. 2022–2032
2032
Closest in time.