Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) makes an agent learn from trial-and-error experiences gathered during the interaction with the environment.
J. A. Hartigan and M. A. Wong, “Algorithm as 136: A k-means clustering algorithm,” Journal of the royal statistical society. series c (applied statistics)
1979
Earlier work this paper cites.
Proceedings of the Multivariate Statistical Workshop for Geologists and Geochemists
S. Wold, K. Esbensen, and P. Geladi, “Principal component analysis,” Chemometrics and Intelligent Laboratory Systems · 1987
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition
2009
Earlier work this paper cites.
T. Hastie, R. Tibshirani, J. Friedman, T. Hastie, R. Tibshirani, and J. Friedman, “Overview of supervised learning,” The elements of statistical learning: Data mining, inference, and prediction
2009
Earlier work this paper cites.
S. Lange, T. Gabel, and M. Riedmiller, “Batch reinforcement learning,” in Reinforcement learning
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” International Conference on Learning Representations
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. E. Sallab, M. Abdou, E. Perot, et al
2017
Earlier work this paper cites.
S. Gu, E. Holly, T. Lillicrap, et al
2017
Earlier work this paper cites.
H. Spieker, A. Gotlieb, D. Marijan, and M. Mossige, “Reinforcement learning for automatic test case prioritization and selection in continuous integration,” in Proceedings of the 26th ACM SIGSOFT International Symposium on Software Testing and Analysis
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “Carla: An open urban driving simulator,” in Conference on robot learning
2017
Earlier work this paper cites.
H. Spieker, A. Gotlieb, D. Marijan, and M. Mossige, “Reinforcement learning for automatic test case prioritization and selection in continuous integration,” in Proceedings of the 26th ACM SIGSOFT International Symposium on Software Testing and Analysis
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver, “Rainbow: Combining improvements in deep reinforcement learning,” in Thirty-second AAAI conference on artificial intelligence
2018
Earlier work this paper cites.
B. Tran, J. Li, and A. Madry, “Spectral signatures in backdoor attacks,” in Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada
2018
Earlier work this paper cites.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Earlier work this paper cites.
S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International Conference on Machine Learning
2018
Earlier work this paper cites.
MIT press, 2018
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction · 2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning
2018
Earlier work this paper cites.
D. Adamo, M. K. Khan, S. Koppula, and R. Bryce, “Reinforcement learning for android gui testing,” in Proceedings of the 9th ACM SIGSOFT International Workshop on Automating TEST Case Design, Selection, and Evaluation
2018
Earlier work this paper cites.
T. Bansal, J. Pachocki, S. Sidor, I. Sutskever, and I. Mordatch, “Emergent complexity via multi-agent competition,” in International Conference on Learning Representations
2018
Earlier work this paper cites.
R. Gupta, A. Kanade, and S. Shevade, “Deep reinforcement learning for syntactic error repair in student programs,” in Proceedings of the AAAI Conference on Artificial Intelligence
2019
Earlier work this paper cites.
O. Nachum, B. Dai, I. Kostrikov, et al
2019
Earlier work this paper cites.
B. Chen, W. Carvalho, N. Baracaldo, H. Ludwig, B. Edwards, T. Lee, I. M. Molloy, and B. Srivastava, “Detecting backdoor attacks on deep neural networks by activation clustering,” in Workshop on Artificial Intelligence Safety 2019 co-located with the Thirty-Third AAAI Conference on Artificial Intelligence 2019 (AAAI-19)
2019
Earlier work this paper cites.
B. Wang, Y. Yao, S. Shan, H. Li, B. Viswanath, H. Zheng, and B. Y. Zhao, “Neural cleanse: Identifying and mitigating backdoor attacks in neural networks,” in 2019 IEEE Symposium on Security and Privacy (SP)
2019
Earlier work this paper cites.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in International Conference on Machine Learning
2019
Earlier work this paper cites.
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine, “Stabilizing off-policy q-learning via bootstrapping error reduction,” Advances in Neural Information Processing Systems
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Zheng, X. Xie, T. Su, L. Ma, J. Hao, Z. Meng, Y. Liu, R. Shen, Y. Chen, and C. Fan, “Wuji: Automatic online combat game testing using evolutionary deep reinforcement learning,” in 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE)
2019
Cited alongside, same era.
Y. Yao, H. Li, H. Zheng, and B. Y. Zhao, “Latent backdoor attacks on deep neural networks,” in Proceedings of the ACM SIGSAC Conference on Computer and Communications Security
2019
Cited alongside, same era.
C. Gong, Y. Bai, X. Hou, and X. Ji, “Stable training of bellman error in reinforcement learning,” in International Conference on Neural Information Processing
2020
Cited alongside, same era.
2020
Cited alongside, same era.
M. H. Asyrofi, Z. Yang, I. N. B. Yusuf, H. J. Kang, F. Thung, and D. Lo, “Biasfinder: Metamorphic test generation to uncover bias for sentiment analysis systems,” IEEE Transactions on Software Engineering
2021
Later among the works it cites.
M. H. Asyrofi, Z. Yang, and D. Lo, “Crossasr++: A modular differential testing framework for automatic speech recognition,” in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering
2021
Later among the works it cites.
Y. Zheng, Y. Liu, X. Xie, Y. Liu, L. Ma, J. Hao, and Y. Liu, “Automatic web testing using curiosity-driven reinforcement learning,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE)
2021
Later among the works it cites.
U. C. Türker, R. M. Hierons, M. R. Mousavi, and I. Y. Tyukin, “Efficient state synchronisation in model-based testing through reinforcement learning,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Yan, S. Mahmud, H. Shen, N. Z. Foutz, and J. Anton, “Mobirescue: Reinforcement learning based rescue team dispatching in a flooding disaster,” in 2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
P. Kiourti, K. Wardega, S. Jha, and W. Li, “Trojdrl: evaluation of backdoor attacks on deep reinforcement learning,” in 2020 57th ACM/IEEE Design Automation Conference (DAC)
2020
Cited alongside, same era.
S. Wang, S. Nepal, C. Rudolph, M. Grobler, S. Chen, and T. Chen, “Backdoor attacks against transfer learning with pre-trained deep learning models,” IEEE Transactions on Services Computing
2020
Cited alongside, same era.
2020
Cited alongside, same era.
S. Zhao, X. Ma, X. Zheng, J. Bailey, J. Chen, and Y.-G. Jiang, “Clean-label backdoor attacks on video recognition models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2020
Cited alongside, same era.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A pre-trained model for programming and natural languages,” in Findings of EMNLP 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
X. Chen, A. Salem, M. Backes, S. Ma, and Y. Zhang, “Badnl: Backdoor attacks against nlp models,” in ICML 2021 Workshop on Adversarial Machine Learning
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Li, Y. Li, B. Wu, L. Li, R. He, and S. Lyu, “Invisible backdoor attack with sample-specific triggers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision
2021
Later among the works it cites.
E. Wenger, J. Passananti, A. N. Bhagoji, Y. Yao, H. Zheng, and B. Y. Zhao, “Backdoor attacks against deep learning systems in the physical world,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2021
Later among the works it cites.
C. Wu, A. R. Kreidieh, K. Parvate, E. Vinitsky, and A. M. Bayen, “Flow: A modular learning framework for mixed autonomy traffic,” IEEE Transactions on Robotics
2021
Later among the works it cites.
A. Levine and S. Feizi, “Deep partition aggregation: Provable defenses against general poisoning attacks,” in International Conference on Learning Representations
2021
Later among the works it cites.
W. Guo, X. Wu, U. Khan, and X. Xing, “Edge: Explaining deep reinforcement learning policies,” in Advances in Neural Information Processing Systems
2021
Later among the works it cites.
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot, “Machine unlearning,” in 2021 IEEE Symposium on Security and Privacy (SP)
2021
Later among the works it cites.
2022
Closest in time.
I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” in International Conference on Learning Representations
2022
Closest in time.
J. M. Zhang, M. Harman, L. Ma, and Y. Liu, “Machine learning testing: Survey, landscapes and horizons,” IEEE Transactions on Software Engineering
2022
Closest in time.
J. M. Zhang, M. Harman, L. Ma, and Y. Liu, “Machine learning testing: Survey, landscapes and horizons,” IEEE Transactions on Software Engineering
2022
Closest in time.
J. Chen, X. Wang, Y. Zhang, H. Zheng, S. Yu, and L. Bao, “Agent manipulator: Stealthy strategy attacks on deep reinforcement learning,” Applied Intelligence
2022
Closest in time.
2022
Closest in time.
Y. Chen, Z. Zheng, and X. Gong, “Marnet: Backdoor attacks against cooperative multi-agent reinforcement learning,” IEEE Transactions on Dependable and Secure Computing
2022
Closest in time.
W. Wang, A. J. Levine, and S. Feizi, “Improved certified defenses against data poisoning with (Deterministic) finite aggregation,” in Proceedings of the 39th International Conference on Machine Learning
2022
Closest in time.
R. Chen, Z. Li, J. Li, J. Yan, and C. Wu, “On collective robustness of bagging against data poisoning,” in Proceedings of the 39th International Conference on Machine Learning
2022
Closest in time.
W. Jiang, X. Wen, J. Zhan, X. Wang, and Z. Song, “Interpretability-guided defense against backdoor attacks to deep neural networks,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
2022
Closest in time.
S. Fang and A. Choromanska, “Backdoor attacks on the dnn interpretation system,” Proceedings of the AAAI Conference on Artificial Intelligence
2022
Closest in time.
Y. Liu, G. Shen, G. Tao, Z. Wang, S. Ma, and X. Zhang, “Complex backdoor detection by symmetric feature differencing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
Closest in time.
Y. Liu, M. Fan, C. Chen, X. Liu, Z. Ma, L. Wang, and J. Ma, “Backdoor defense with machine unlearning,” in IEEE INFOCOM - IEEE Conference on Computer Communications
2022
Closest in time.
Y. Zeng, S. Chen, W. Park, Z. Mao, M. Jin, and R. Jia, “Adversarial unlearning of backdoors via implicit hypergradient,” in International Conference on Learning Representations
2022
Closest in time.
S. Bharti, X. Zhang, A. Singla, and J. Zhu, “Provable defense against backdoor policies in reinforcement learning,” Advances in Neural Information Processing Systems
2022
Closest in time.
T. Seno and M. Imai, “D3rlpy: An offline deep reinforcement learning library,” J. Mach. Learn. Res
2023
Closest in time.
2023
Closest in time.
Y. Mao, Z. Xin, Z. Li, J. Hong, Q. Yang, and S. Zhong, “Secure split learning against property inference, data reconstruction, and feature space hijacking attacks,” in European Symposium on Research in Computer Security
2023
Closest in time.