Fetching the paper…
Reading the bibliography…
A deep reinforcement learning (DRL) agent observes its states through observations, which may contain natural measurement errors or adversarial noises.
Recursive stochastic algorithms for global optimization in ℝ d \mathbb{R}^{d}
Gelfand, S. B. and Mitter, S. K · 1991
Earlier work this paper cites.
Artificial life and real robots
Brooks, R. A · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Tan, M · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Littman, M. L · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Robustness in Markov decision problems with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2004
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Bu, L., Babu, R., De Schutter, B., et al · 2008
Earlier work this paper cites.
Double q-learning
Hasselt, H. V · 2010
Earlier work this paper cites.
Distributionally robust markov decision processes
Xu, H. and Mannor, S · 2010
Earlier work this paper cites.
Safe policy iteration
Pirotta, M., Restelli, M., Pecorino, A., and Calandriello, D · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Finite-time analysis of projected Langevin Monte Carlo
Bubeck, S., Eldan, R., and Lehec, J · 2015
Earlier work this paper cites.
Trust region policy optimization
Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Distributional smoothing with virtual adversarial training
Miyato, T., Maeda, S.-i., Koyama, M., Nakae, K., and Ishii, S · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Robust partially observable Markov decision process
Osogami, T · 2015
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Continuous deep Q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Earlier work this paper cites.
Adversarial machine learning at scale
Kurakin, A., Goodfellow, I., and Bengio, S · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Deep reinforcement learning with double Q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S., Papernot, N., Goodfellow, I., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Fast and effective robustness certification
Singh, G., Gehr, T., Mirman, M., Püschel, M., and Vechev, M · 2018
Later among the works it cites.
Towards fast computation of certified robustness for ReLU networks
Weng, T.-W., Zhang, H., Chen, H., Song, Z., Hsieh, C.-J., Daniel, L., Boning, D., and Dhillon, I · 2018
Later among the works it cites.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Wong, E. and Kolter, Z · 2018
Later among the works it cites.
Scaling provable adversarial defenses
Wong, E., Schmidt, F., Metzen, J. H., and Kolter, J. Z · 2018
Later among the works it cites.
Global convergence of langevin dynamics based algorithms for nonconvex optimization
Xu, P., Chen, J., Zou, D., and Gu, Q · 2018
Later among the works it cites.
Efficient neural network robustness certification with general activation functions
Zhang, H., Weng, T.-W., Chen, P.-Y., Hsieh, C.-J., and Daniel, L · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Delving into adversarial attacks on deep policies
Kos, J. and Song, D · 2017
Cited alongside, same era.
Tactics of adversarial attack on deep reinforcement learning agents
Lin, Y.-C., Hong, Z.-W., Liao, Y.-H., Shih, M.-L., Liu, M.-Y., and Sun, M · 2017
Cited alongside, same era.
Adversarially robust policy learning: Active construction of physically-plausible perturbations
Mandlekar, A., Zhu, Y., Garg, A., Fei-Fei, L., and Savarese, S · 2017
Cited alongside, same era.
Virtual to real reinforcement learning for autonomous driving
Pan, X., You, Y., Wang, Z., and Lu, C · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient Langevin dynamics: a nonasymptotic analysis
Raginsky, M., Rakhlin, A., and Telgarsky, M · 2017
Cited alongside, same era.
Deep reinforcement learning framework for autonomous driving
Sallab, A. E., Abdou, M., Perot, E., and Yogamani, S · 2017
Cited alongside, same era.
Later among the works it cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Later among the works it cites.
Adversarial training and provable defenses: Bridging the gap
Balunovic, M. and Vechev, M · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dębiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Later among the works it cites.
Online robustness training for deep reinforcement learning
Fischer, M., Mirman, M., and Vechev, M · 2019
Later among the works it cites.
Adversary A3C for robust reinforcement learning
Gu, Z., Jia, Z., and Choset, H · 2019
Later among the works it cites.
Robust multi-agent reinforcement learning via minimax deep deterministic policy gradient
Li, S., Wu, Y., Cui, X., Dong, H., Fang, F., and Russell, S · 2019
Later among the works it cites.
Certified adversarial robustness for deep reinforcement learning
Lütjens, B., Everett, M., and How, J. P · 2019
Later among the works it cites.
Robust reinforcement learning for continuous control with model misspecification
Mankowitz, D. J., Levine, N., Jeong, R., Abdolmaleki, A., Springenberg, J. T., Mann, T., Hester, T., and Riedmiller, M · 2019
Later among the works it cites.
Optimal attacks on reinforcement learning policies
Russo, A. and Proutiere, A · 2019
Later among the works it cites.
A convex relaxation barrier to tight robustness verification of neural networks
Salman, H., Yang, G., Zhang, H., Hsieh, C.-J., and Zhang, P · 2019
Later among the works it cites.
An abstract domain for certifying neural networks
Singh, G., Gehr, T., Püschel, M., and Vechev, M · 2019
Later among the works it cites.
Action robust reinforcement learning and applications in continuous control
Tessler, C., Efroni, Y., and Mannor, S · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Introducing voyage deepdrive -unlocking the potential of deep reinforcement learning
Voyage · 2019
Later among the works it cites.
Characterizing attacks on deep reinforcement learning
Xiao, C., Pan, X., He, W., Peng, J., Sun, M., Yi, J., Li, B., and Song, D · 2019
Later among the works it cites.
Advanced planning for autonomous vehicles using reinforcement learning and deep inverse reinforcement learning
You, C., Lu, J., Filev, D., and Tsiotras, P · 2019
Later among the works it cites.
Theoretically principled trade-off between robustness and accuracy
Zhang, H., Yu, Y., Jiao, J., Xing, E. P., Ghaoui, L. E., and Jordan, M. I · 2019
Later among the works it cites.
Implementation matters in deep policy gradients: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Closest in time.
Challenges and countermeasures for adversarial attacks on deep reinforcement learning
Ilahi, I., Usama, M., Qadir, J., Janjua, M. U., Al-Fuqaha, A., Hoang, D. T., and Niyato, D · 2020
Closest in time.
Deep reinforcement learning with smooth policy
Shen, Q., Li, Y., Jiang, H., Wang, Z., and Zhao, T · 2020
Closest in time.
Automatic perturbation analysis on general computational graphs
Xu, K., Shi, Z., Zhang, H., Huang, M., Chang, K.-W., Kailkhura, B., Lin, X., and Hsieh, C.-J · 2020
Closest in time.
Towards stable and efficient training of verifiably robust neural networks
Zhang, H., Chen, H., Xiao, C., Li, B., Boning, D., and Hsieh, C.-J · 2020
Closest in time.