Fetching the paper…
Reading the bibliography…
Deep Reinforcement Learning (DRL) has made significant advancements in various fields, such as autonomous driving, healthcare, and robotics, by enabling agents to learn optimal policies through interactions with their environments.
A. W. Moore, “Efficient memory-based learning for robot control,” University of Cambridge, Computer Laboratory, Tech. Rep., 1990
1990
Earlier work this paper cites.
V. J. Easton and J. H. McColl, “Statistics glossary v1. 1,” 1997
1997
Earlier work this paper cites.
K. Hansen, A. Ravn, and V. Stavridou, “From safety analysis to software requirements,” IEEE Transactions on Software Engineering , vol. 24, no. 7, pp. 573–584, 1998
1998
Earlier work this paper cites.
K. Doya, “Reinforcement learning in continuous time and space,” Neural computation , vol. 12, no. 1, pp. 219–245, 2000
2000
Earlier work this paper cites.
L. Breiman, “Random forests,” Mach. Learn. , vol. 45, no. 1, p. 5–32, Oct. 2001. [Online]. Available: https://doi.org/10.1023/A:1010933404324
2001
Earlier work this paper cites.
J. H. Friedman, “Greedy function approximation: a gradient boosting machine,” Annals of statistics , pp. 1189–1232, 2001
2001
Earlier work this paper cites.
L. Li, T. J. Walsh, and M. L. Littman, “Towards a unified theory of state abstraction for mdps.” ISAIM , vol. 4, no. 5, p. 9, 2006
2006
Earlier work this paper cites.
R. Caruana and A. Niculescu-Mizil, “An empirical comparison of supervised learning algorithms,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 161–168
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
J. Schaeffer, N. Sturtevant, R. Holte, and K. Anderson, “Coarse-to-fine search techniques,” 2008
2008
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. Pecka and T. Svoboda, “Safe exploration techniques for reinforcement learning–an overview,” in Modelling and Simulation for Autonomous Systems: First International Workshop, MESAS 2014, Rome, Italy, May 5-6, 2014, Revised Selected Papers 1 . Springer, 2014, pp. 357–375
2014
Earlier work this paper cites.
H. van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” 2015
2015
Earlier work this paper cites.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015
2015
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 30, no. 1, 2016
2016
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas, “Dueling network architectures for deep reinforcement learning,” 2016
2016
Earlier work this paper cites.
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Processing Magazine , vol. 34, no. 6, pp. 26–38, 2017
2017
Earlier work this paper cites.
G. Katz, C. Barrett, D. L. Dill, K. Julian, and M. J. Kochenderfer, “Reluplex: An efficient smt solver for verifying deep neural networks,” in Computer Aided Verification: 29th International Conference, CAV 2017, Heidelberg, Germany, July 24-28, 2017, Proceedings, Part I 30 . Springer, 2017, pp. 97–117
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu, “Safe reinforcement learning via shielding,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
“Road vehicles – functional safety,” ISO 26262:2018, 2018
2018
Earlier work this paper cites.
D. Abel, D. Arumugam, L. Lehnert, and M. Littman, “State abstractions for lifelong reinforcement learning,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, J. Dy and A. Krause, Eds., vol. 80. PMLR, 10–15 Jul 2018, pp. 10–19. [Online]. Available: http://proceedings.mlr.press/v80/abel18a.html
2018
Earlier work this paper cites.
N. Jiang, “Notes on state abstractions,” 2018
2018
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
R. Akrour, F. Veiga, J. Peters, and G. Neumann, “Regularizing reinforcement learning with state abstraction,” in 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2018, pp. 534–539
2018
Earlier work this paper cites.
A. Tavakoli, F. Pardo, and P. Kormushev, “Action branching architectures for deep reinforcement learning,” in Proceedings of the aaai conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
E. Leurent, “An environment for autonomous driving decision-making,” https://github.com/eleurent/highway-env , 2018
2018
Earlier work this paper cites.
A. Hill, A. Raffin, M. Ernestus, A. Gleave, A. Kanervisto, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu, “Stable baselines,” https://github.com/hill-a/stable-baselines , 2018
2018
Earlier work this paper cites.
R. Michelmore, M. Kwiatkowska, and Y. Gal, “Evaluating uncertainty quantification in end-to-end autonomous driving control,” arXiv e-prints , pp. arXiv–1811, 2018
2018
Earlier work this paper cites.
E. Bartocci, J. Deshmukh, A. Donzé, G. Fainekos, O. Maler, D. Ničković, and S. Sankaranarayanan, “Specification-based monitoring of cyber-physical systems: a survey on theory, tools and applications,” Lectures on Runtime Verification: Introductory and Advanced Topics , pp. 135–175, 2018
2018
Earlier work this paper cites.
M. Strickland, G. Fainekos, and H. B. Amor, “Deep predictive models for collision risk assessment in autonomous driving,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 4685–4692
2018
Cited alongside, same era.
“Road vehicles — safety of the intended functionality,” ISO/PAS 21448:2019, 2019
2019
Cited alongside, same era.
G. Dulac-Arnold, D. Mankowitz, and T. Hester, “Challenges of real-world reinforcement learning,” 2019
2019
Cited alongside, same era.
B. Jang, M. Kim, G. Harerimana, and J. W. Kim, “Q-learning algorithms: A comprehensive classification and applications,” IEEE access , vol. 7, pp. 133 653–133 667, 2019
2019
Cited alongside, same era.
2022
Later among the works it cites.
B. Könighofer, J. Rudolf, A. Palmisano, M. Tappler, and R. Bloem, “Online shielding for reinforcement learning,” Innovations in Systems and Software Engineering , pp. 1–16, 2022
2022
Later among the works it cites.
M. Liu, L. Li, S. Hao, Y. Zhu, and D. Zhao, “Soft contrastive learning with q-irrelevance abstraction for reinforcement learning,” IEEE Transactions on Cognitive and Developmental Systems , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
P. Mallozzi, E. Castellano, P. Pelliccione, G. Schneider, and K. Tei, “A runtime monitoring framework to enforce invariants on reinforcement learning agents exploring complex environments,” in 2019 IEEE/ACM 2nd International Workshop on Robotics Software Engineering (RoSE) . IEEE, 2019, pp. 5–12
2019
Cited alongside, same era.
C.-H. Cheng, G. Nührenberg, and H. Yasuoka, “Runtime monitoring neuron activation patterns,” in 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE) . IEEE, 2019, pp. 300–303
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Turchetta, A. Kolobov, S. Shah, A. Krause, and A. Agarwal, “Safe reinforcement learning via curriculum induction,” Advances in Neural Information Processing Systems , vol. 33, pp. 12 151–12 162, 2020
2020
Cited alongside, same era.
Y. Tang and S. Agrawal, “Discretizing continuous action space for on-policy optimization,” in Proceedings of the aaai conference on artificial intelligence , vol. 34, no. 04, 2020, pp. 5981–5988
2020
Cited alongside, same era.
D. Steckelmacher, H. Plisnier, D. M. Roijers, and A. Nowé, “Sample-efficient model-free reinforcement learning with off-policy critics,” in Machine Learning and Knowledge Discovery in Databases , U. Brefeld, E. Fromont, A. Hotho, A. Knobbe, M. Maathuis, and C. Robardet, Eds. Cham: Springer International Publishing, 2020, pp. 19–34
2020
Cited alongside, same era.
“Openai,” https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html , 2018, [Accessed 24 Jan. 2022.]
2022
Later among the works it cites.
2022
Later among the works it cites.
D. Melcer, C. Amato, and S. Tripakis, “Shield decentralization for safe multi-agent reinforcement learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 13 367–13 379, 2022
2022
Later among the works it cites.
A. Stocco, P. J. Nunes, M. D’Amorim, and P. Tonella, “Thirdeye: Attention maps for safe autonomous driving systems,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering , 2022, pp. 1–12
2022
Later among the works it cites.
M. Hussain, N. Ali, and J.-E. Hong, “Deepguard: A framework for safeguarding autonomous driving systems from inconsistent behaviour,” Automated Software Engineering , vol. 29, no. 1, p. 1, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Zolfagharian, M. Abdellatif, L. C. Briand, M. Bagherzadeh, and S. Ramesh, “A search-based testing approach for deep reinforcement learning agents,” IEEE Transactions on Software Engineering , 2023
2023
Closest in time.
E. Marchesini, L. Marzari, A. Farinelli, and C. Amato, “Safe deep reinforcement learning by verifying task-level properties,” in Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems , 2023, pp. 1466–1475
2023
Closest in time.
S. Antony, R. Roy, and Y. Bi, “Q-learning: Solutions for grid world problem with forward and backward reward propagations,” in Artificial Intelligence XL , M. Bramer and F. Stahl, Eds. Cham: Springer Nature Switzerland, 2023, pp. 266–271
2023
Closest in time.
Z. Aghababaeyan, M. Abdellatif, L. Briand, S. Ramesh, and M. Bagherzadeh, “Black-box testing of deep neural networks through test case diversity,” IEEE Transactions on Software Engineering , 2023
2023
Closest in time.
Y. Okawa, T. Sasaki, H. Yanami, and T. Namerikawa, “Safe exploration method for reinforcement learning under existence of disturbance,” 2023
2023
Closest in time.
F. Bellotti, L. Lazzaroni, A. Capello, M. Cossu, A. De Gloria, and R. Berta, “Explaining a deep reinforcement learning (drl)-based automated driving agent in highway simulations,” IEEE Access , vol. 11, pp. 28 522–28 550, 2023
2023
Closest in time.
A. Beikmohammadi and S. Magnússon, “Comparing nars and reinforcement learning: An analysis of ona and q-learning algorithms,” in Artificial General Intelligence , P. Hammer, M. Alirezaie, and C. Strannegård, Eds. Cham: Springer Nature Switzerland, 2023, pp. 21–31
2023
Closest in time.
S. Feng, H. Sun, X. Yan, H. Zhu, Z. Zou, S. Shen, and H. X. Liu, “Dense reinforcement learning for safety validation of autonomous vehicles,” Nature , vol. 615, no. 7953, pp. 620–627, 2023
2023
Closest in time.
2023
Closest in time.
J. Cao, Z. Guo, Y. Lv, M. Xu, C. Huang, and H. Liang, “Pollution risk prediction for cadmium in soil from an abandoned mine based on random forest model,” International Journal of Environmental Research and Public Health , vol. 20, no. 6, p. 5097, Mar 2023. [Online]. Available: http://dx.doi.org/10.3390/ijerph20065097
2023
Closest in time.
F. Semeraro, A. Griffiths, and A. Cangelosi, “Human–robot collaboration and machine learning: A systematic review of recent research,” Robotics and Computer-Integrated Manufacturing , vol. 79, p. 102432, 2023
2023
Closest in time.
X. Xie, J. Song, Z. Zhou, F. Zhang, and L. Ma, “Mosaic: Model-based safety analysis framework for ai-enabled cyber-physical systems,” 2023
2023
Closest in time.
2024
Closest in time.
Z. Aghababaeyan, M. Abdellatif, M. Dadkhah, and L. Briand, “Deepgd: A multi-objective black-box test selection approach for deep neural networks,” 2024
2024
Closest in time.
2024
Closest in time.
L. Wen, E. H. Tseng, H. Peng, and S. Zhang, “Dream to adapt: Meta reinforcement learning by latent context imagination and mdp imagination,” IEEE Robotics and Automation Letters , pp. 1–8, 2024
2024
Closest in time.
A. Pighetti, F. Bellotti, C. Oh, L. Lazzaroni, L. Forneris, M. Fresta, and R. Berta, “Investigating adversarial policy learning for robust agents in automated driving highway simulations,” in Applications in Electronics Pervading Industry, Environment and Society , F. Bellotti, M. D. Grammatikakis, A. Mansour, M. Ruo Roch, R. Seepold, A. Solanas, and R. Berta, Eds. Cham: Springer Nature Switzerland, 2024, pp. 124–129
2024
Closest in time.
“Replication package,” https://github.com/amirhosseinzlf/SMARLA , (accessed: 22.10.2024)
2024
Closest in time.
A. Pattanaik, Z. Tang, S. Liu, G. Bommannan, and G. Chowdhary, “Robust deep reinforcement learning with adversarial attacks,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems , 2018, pp. 2040–2042
2042
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in International conference on machine learning . PMLR, 2019, pp. 2052–2062
2062
Closest in time.