Fetching the paper…
Reading the bibliography…
We study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker.
Automatic programming of behavior-based robots using reinforcement learning
Sridhar Mahadevan and Jonathan Connell · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
H-learning: A reinforcement learning method to optimize undiscounted average reward
Prasad Tadepalli and DoKyeong Ok · 1994
Earlier work this paper cites.
On the complexity of teaching
Sally A Goldman and Michael J Kearns · 1995
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Sridhar Mahadevan · 1996
Earlier work this paper cites.
Essentials of stochastic processes , volume 1
Richard Durrett and R Durrett · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Learning rates for q-learning
Eyal Even-Dar and Yishay Mansour · 2003
Earlier work this paper cites.
Experts in a markov decision process
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
Pac model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
Interactive robot task training through dialog and demonstration
Paul E Rybski, Kevin Yoon, Jeremy Stolarz, and Manuela M Veloso · 2007
Earlier work this paper cites.
Potential-based shaping in model-based reinforcement learning
John Asmuth, Michael L Littman, and Robert Zinkov · 2008
Earlier work this paper cites.
Value-based policy teaching with active indirect elicitation
Haoqi Zhang and David C. Parkes · 2008
Earlier work this paper cites.
Policy teaching through reward function learning
Haoqi Zhang, David C. Parkes, and Yiling Chen · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
Adversarial machine learning
Ling Huang, Anthony D Joseph, Blaine Nelson, Benjamin IP Rubinstein, and J Doug Tygar · 2011
Earlier work this paper cites.
What’s clicking what? techniques and innovations of today’s clickbots
Brad Miller, Paul Pearce, Chris Grier, Christian Kreibich, and Vern Paxson · 2011
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
Algorithmic and human teaching of sequential decision tasks
Maya Cakmak and Manuel Lopes · 2012
Earlier work this paper cites.
Dynamic teaching in sequential decision making environments
Thomas J. Walsh and Sergiu Goschin · 2012
Cited alongside, same era.
On actively teaching the crowd to classify
Adish Singla, Ilija Bogunovic, G Bartók, A Karbasi, and A Krause · 2013
Cited alongside, same era.
Simple and scalable response prediction for display advertising
Olivier Chapelle, Eren Manavoglu, and Rómer Rosales · 2014
Cited alongside, same era.
Near-optimally teaching the crowd to classify
Adish Singla, Ilija Bogunovic, Gábor Bartók, Amin Karbasi, and Andreas Krause · 2014
Cited alongside, same era.
Using machine teaching to identify optimal training-set attacks on machine learners
Shike Mei and Xiaojin Zhu · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Data poisoning attacks in contextual bandits
Yuzhe Ma, Kwang-Sung Jun, Lihong Li, and Xiaojin Zhu · 2018
Later among the works it cites.
An algorithmic perspective on imitation learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J Andrew Bagnell, Pieter Abbeel, Jan Peters, et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Sequential attacks on agents for long-term adversarial goals
Edgar Tretschk, Seong Joon Oh, and Mario Fritz · 2018
Later among the works it cites.
An optimal control view of adversarial machine learning
Xiaojin Zhu · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Is feature selection secure against training data poisoning?
Huang Xiao, Battista Biggio, Gavin Brown, Giorgio Fumera, Claudia Eckert, and Fabio Roli · 2015
Cited alongside, same era.
Machine teaching: An inverse problem to machine learning and an approach toward optimal education
Xiaojin Zhu · 2015
Cited alongside, same era.
Data poisoning attacks against autoregressive models
Scott Alfeld, Xiaojin Zhu, and Paul Barford · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Cited alongside, same era.
Data poisoning attacks on factorization-based collaborative filtering
Bo Li, Yining Wang, Aarti Singh, and Yevgeniy Vorobeychik · 2016
Cited alongside, same era.
Xiaojin Zhu, Adish Singla, Sandra Zilles, and Anna N Rafferty · 2018
Later among the works it cites.
Machine teaching for inverse reinforcement learning: Algorithms and applications
Daniel S Brown and Scott Niekum · 2019
Later among the works it cites.
Adversarial attack and defense in reinforcement learning from AI security view
Tong Chen, Jiqiang Liu, Yingxiao Xiang, Wenjia Niu, Endong Tong, and Zhen Han · 2019
Later among the works it cites.
A survey on transfer learning for multiagent reinforcement learning systems
Felipe Leno Da Silva and Anna Helena Reali Costa · 2019
Later among the works it cites.
Deceptive reinforcement learning under adversarial manipulations on cost signals
Yunhan Huang and Quanyan Zhu · 2019
Later among the works it cites.
Interactive teaching algorithms for inverse reinforcement learning
Parameswaran Kamalaruban, Rati Devidze, Volkan Cevher, and Adish Singla · 2019
Later among the works it cites.
Reinforcement Learning for Cyber-Physical Systems: with Cybersecurity Case Studies
Chong Li and Meikang Qiu · 2019
Later among the works it cites.
Data poisoning attacks on stochastic bandits
Fang Liu and Ness B. Shroff · 2019
Later among the works it cites.
Policy poisoning in batch reinforcement learning and control
Yuzhe Ma, Xuezhou Zhang, Wen Sun, and Jerry Zhu · 2019
Later among the works it cites.
Preference-based batch and sequential teaching: Towards a unified view of models
Farnam Mansouri, Yuxin Chen, Ara Vartanian, Jerry Zhu, and Adish Singla · 2019
Later among the works it cites.
Machine teaching of active sequential learners
Tomi Peltola, Mustafa Mert Çelikok, Pedram Daee, and Samuel Kaski · 2019
Later among the works it cites.
Learner-aware teaching: Inverse reinforcement learning with preferences and constraints
Sebastian Tschiatschek, Ahana Ghosh, Luis Haug, Rati Devidze, and Adish Singla · 2019
Later among the works it cites.
Understanding the power and limitations of teaching with imperfect knowledge
Rati Devidze, Farnam Mansouri, Luis Haug, Yuxin Chen, and Adish Singla · 2020
Closest in time.
Policy teaching via environment poisoning: Training-time adversarial attacks against reinforcement learning
Amin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu, and Adish Singla · 2020
Closest in time.
Adaptive reward-poisoning attacks against reinforcement learning
Xuezhou Zhang, Yuzhe Ma, Adish Singla, and Xiaojin Zhu · 2020
Closest in time.