Fetching the paper…
Reading the bibliography…
While real-world applications of reinforcement learning are becoming popular, the security and robustness of RL systems are worthy of more attention and exploration.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
A reinforcement learning framework for parameter control in computer vision applications
Graham W Taylor · 2004
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Detecting outliers: Do not use standard deviation around the mean, use absolute deviation around the median
Christophe Leys, Christophe Ley, Olivier Klein, Philippe Bernard, and Laurent Licata · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Earlier work this paper cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2017
Earlier work this paper cites.
Carla: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne · 2017
Earlier work this paper cites.
Trojaning attack on neural networks
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang · 2017
Earlier work this paper cites.
Jpmorgan develops robot to execute trades
Laura Noonan · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Reinforcement learning in computer vision
AV Bernstein and Evgeny V Burnaev · 2018
Earlier work this paper cites.
Adversarial policies: Attacking deep reinforcement learning
Adam Gleave, Michael Dennis, Cody Wild, Neel Kant, Sergey Levine, and Stuart Russell · 2019
Earlier work this paper cites.
Tabor: A highly accurate approach to inspecting and restoring trojan backdoors in ai systems
Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song · 2019
Earlier work this paper cites.
Abs: Scanning neural networks for back-doors by artificial brain stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang · 2019
Earlier work this paper cites.
Solving rubik’s cube with a robot hand, 2019
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 2019
Earlier work this paper cites.
Stable baselines3
Antonin Raffin, Ashley Hill, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, and Noah Dormann · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander S. Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom L. Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver · 2019
Cited alongside, same era.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Cited alongside, same era.
Latent backdoor attacks on deep neural networks
Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y Zhao · 2019
Cited alongside, same era.
Trojdrl: evaluation of backdoor attacks on deep reinforcement learning
Panagiota Kiourti, Kacper Wardega, Susmit Jha, and Wenchao Li · 2020
Cited alongside, same era.
Deep reinforcement learning for traffic signal control: A review
Mind your data! hiding backdoors in offline reinforcement learning datasets
Chen Gong, Zhou Yang, Yunpeng Bai, Junda He, Jieke Shi, Arunesh Sinha, Bowen Xu, Xinwen Hou, Guoliang Fan, and David Lo · 2022
Closest in time.
AEVA: Black-box backdoor detection using adversarial extreme value analysis
Junfeng Guo, Ang Li, and Cong Liu · 2022
Closest in time.
Backdoor defense via decoupling the training process
Kunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin, and Kui Ren · 2022
Closest in time.
Deep reinforcement learning in computer vision: a comprehensive survey
Ngan Le, Vidhiwar Singh Rathour, Kashu Yamazaki, Khoa Luu, and Marios Savvides · 2022
Closest in time.
Backdoor learning: A survey
Yiming Li, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2022
Closest in time.
Retrievalguard: Provably robust 1-nearest neighbor image retrieval
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Faizan Rasheed, Kok-Lim Alvin Yau, Rafidah Md. Noor, Celimuge Wu, and Yeh-Ching Low · 2020
Cited alongside, same era.
Practical detection of trojan neural networks: Data-limited and data-free cases
Ren Wang, Gaoyuan Zhang, Sijia Liu, Pin-Yu Chen, Jinjun Xiong, and Meng Wang · 2020
Cited alongside, same era.
Black-box detection of backdoor attacks with limited information and data
Yinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang, Zihao Xiao, Hang Su, and Jun Zhu · 2021
Cited alongside, same era.
Adversarial policy learning in two-player competitive games
Wenbo Guo, Xian Wu, Sui Huang, and Xinyu Xing · 2021
Cited alongside, same era.
Edge: Explaining deep reinforcement learning policies
Wenbo Guo, Xian Wu, Usmann Khan, and Xinyu Xing · 2021
Cited alongside, same era.
Invisible backdoor attack with sample-specific triggers
Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu · 2021
Cited alongside, same era.
Backdoor attack in the physical world
Yiming Li, Tongqing Zhai, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2021
Cited alongside, same era.
Backdoor scanning for deep neural networks through k-arm optimization
Guangyu Shen, Yingqi Liu, Guanhong Tao, Shengwei An, Qiuling Xu, Siyuan Cheng, Shiqing Ma, and Xiangyu Zhang · 2021
Cited alongside, same era.
Yihan Wu, Hongyang Zhang, and Heng Huang · 2022
Closest in time.
Toward unified data and algorithm fairness via adversarial data augmentation and adaptive model fine-tuning
Yanfu Zhang, Runxue Bao, Jian Pei, and Heng Huang · 2022
Closest in time.
Recover fair deep classification models via altering pre-trained structure
Yanfu Zhang, Shangqian Gao, and Heng Huang · 2022
Closest in time.
Federated incremental semantic segmentation
Jiahua Dong, Duzhen Zhang, Yang Cong, Wei Cong, Henghui Ding, and Dengxin Dai · 2023
Closest in time.
Learning to jointly share and prune weights for grounding based vision and language models
Shangqian Gao, Burak Uzkent, Yilin Shen, Heng Huang, and Hongxia Jin · 2023
Closest in time.
Not all samples are born equal: Towards effective clean-label backdoor attacks
Yinghua Gao, Yiming Li, Linghui Zhu, Dongxian Wu, Yong Jiang, and Shu-Tao Xia · 2023
Closest in time.
Scale-up: An efficient black-box input-level backdoor detection via analyzing scaled prediction consistency
Junfeng Guo, Yiming Li, Xun Chen, Hanqing Guo, Lichao Sun, and Cong Liu · 2023
Closest in time.
Red: A systematic real-time scheduling approach for robotic environmental dynamics, 2023
Zexin Li, Tao Ren, Xiaoxi He, and Cong Liu · 2023
Closest in time.
R3: On-device real-time deep reinforcement learning for autonomous robotics, 2023
Zexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio, Hyoseung Kim, and Cong Liu · 2023
Closest in time.
Pimbot: Policy and incentive manipulation for multi-robot reinforcement learning in social dilemmas, 2023
Shahab Nikkhoo, Zexin Li, Aritra Samanta, Yufei Li, and Cong Liu · 2023
Closest in time.
Revisiting the assumption of latent separability for backdoor defenses
Xiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar, and Prateek Mittal · 2023
Closest in time.
UNICORN: A unified backdoor trigger inversion framework
Zhenting Wang, Kai Mei, Juan Zhai, and Shiqing Ma · 2023
Closest in time.
Adversarial weight perturbation improves generalization in graph neural networks
Yihan Wu, Aleksandar Bojchevski, and Heng Huang · 2023
Closest in time.
Tdc: Towards extremely efficient cnns on gpus via hardware-aware tucker decomposition
Lizhi Xiang, Miao Yin, Chengming Zhang, Aravind Sukumaran-Rajam, P Sadayappan, Bo Yuan, and Dingwen Tao · 2023
Closest in time.
Comcat: Towards efficient compression and customization of attention-based vision models
Jinqi Xiao, Miao Yin, Yu Gong, Xiao Zang, Jian Ren, and Bo Yuan · 2023
Closest in time.