Fetching the paper…
Reading the bibliography…
We study black-box reward poisoning attacks against reinforcement learning (RL), in which an adversary aims to manipulate the rewards to mislead a sequence of RL agents with unknown algorithms to learn a nefarious policy in an environment unknown to the adversary a priori.
On the complexity of teaching
Goldman, S. and Kearns, M · 1995
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J · 2003
Earlier work this paper cites.
The sample complexity of exploration in the multi-armed bandit problem
Mannor, S. and Tsitsiklis, J. N · 2004
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Potential-based shaping in model-based reinforcement learning
Asmuth, J., Littman, M. L., and Zinkov, R · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Value-based policy teaching with active indirect elicitation
Zhang, H. and Parkes, D. C · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R · 2009
Earlier work this paper cites.
A structured multiarmed bandit problem and the greedy policy
Mersereau, A. J., Rusmevichientong, P., and Tsitsiklis, J. N · 2009
Earlier work this paper cites.
Policy teaching through reward function learning
Zhang, H., Parkes, D. C., and Chen, Y · 2009
Earlier work this paper cites.
Policy teaching in reinforcement learning via environment poisoning attacks
Rakhsha, A., Radanovic, G., Devidze, R., Zhu, X., and Singla, A · 2011
Earlier work this paper cites.
Algorithmic and human teaching of sequential decision tasks
Cakmak, M. and Lopes, M · 2012
Earlier work this paper cites.
Dynamic teaching in sequential decision making environments
Walsh, T. J. and Goschin, S · 2012
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J., Shlens, J., and Szegedy, C · 2014
Earlier work this paper cites.
Near-optimally teaching the crowd to classify
Singla, A., Bogunovic, I., Bartók, G., Karbasi, A., and Krause, A · 2014
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Machine teaching: an inverse problem to machine learning and an approach toward optimal education
Zhu, X · 2015
Cited alongside, same era.
Towards end-to-end reinforcement learning of dialogue agents for information access
Dhingra, B., Li, L., Li, X., Gao, J., Chen, Y.-N., Ahmed, F., and Deng, L · 2016
Cited alongside, same era.
Cooperative inverse reinforcement learning
Hadfield-Menell, D., Russell, S. J., Abbeel, P., and Dragan, A · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Li, J., Monroe, W., Ritter, A., Galley, M., Gao, J., and Jurafsky, D · 2016
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, I. and Van Roy, B · 2016
Cited alongside, same era.
Deep reinforcement learning for page-wise recommendations
Zhao, X., Xia, L., Zhang, L., Ding, Z., Yin, D., and Tang, J · 2018
Later among the works it cites.
An overview of machine teaching
Zhu, X., Singla, A., Zilles, S., and Rafferty, A. N · 2018
Later among the works it cites.
Best arm identification for contaminated bandits
Altschuler, J., Brunel, V.-E., and Malek, A · 2019
Later among the works it cites.
Machine teaching for inverse reinforcement learning: Algorithms and applications
Brown, D. S. and Niekum, S · 2019
Later among the works it cites.
Top-k off-policy correction for a reinforce recommender system
Chen, M., Beutel, A., Covington, P., Jain, S., Belletti, F., and Chi, E. H · 2019
Later among the works it cites.
Teaching a black-box learner
Dasgupta, S., Hsu, D., Poulis, S., and Zhu, X · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Vulnerability of deep reinforcement learning to policy induction attacks
Behzadan, V. and Munir, A · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S., Papernot, N., Goodfellow, I., Duan, Y., and Abbeel, P · 2017
Cited alongside, same era.
Delving into adversarial attacks on deep policies
Kos, J. and Song, D · 2017
Cited alongside, same era.
Tactics of adversarial attack on deep reinforcement learning agents
Lin, Y.-C., Hong, Z.-W., Liao, Y.-H., Shih, M.-L., Liu, M.-Y., and Sun, M · 2017
Cited alongside, same era.
Asymptotic theory of weakly dependent random processes , volume 80
Rio, E · 2017
Cited alongside, same era.
Later among the works it cites.
Deceptive reinforcement learning under adversarial manipulations on cost signals
Huang, Y. and Zhu, Q · 2019
Later among the works it cites.
Interactive teaching algorithms for inverse reinforcement learning
Kamalaruban, P., Devidze, R., Cevher, V., and Singla, A · 2019
Later among the works it cites.
Data poisoning attacks on stochastic bandits
Liu, F. and Shroff, N · 2019
Later among the works it cites.
Policy poisoning in batch reinforcement learning and control
Ma, Y., Zhang, X., Sun, W., and Zhu, J · 2019
Later among the works it cites.
Preference-based batch and sequential teaching: Towards a unified view of models
Mansouri, F., Chen, Y., Vartanian, A., Zhu, J., and Singla, A · 2019
Later among the works it cites.
Machine teaching of active sequential learners
Peltola, T., Çelikok, M. M., Daee, P., and Kaski, S · 2019
Later among the works it cites.
Learner-aware teaching: Inverse reinforcement learning with preferences and constraints
Tschiatschek, S., Ghosh, A., Haug, L., Devidze, R., and Singla, A · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z · 2020
Later among the works it cites.
Teaching with limited information on the learner’s behaviour
Cicalese, F., Filho, S., Laber, E., and Molinaro, M · 2020
Later among the works it cites.
Understanding the power and limitations of teaching with imperfect knowledge
Devidze, R., Mansouri, F., Haug, L., Chen, Y., and Singla, A · 2020
Later among the works it cites.
Reward-free exploration for reinforcement learning
Jin, C., Krishnamurthy, A., Simchowitz, M., and Yu, T · 2020
Later among the works it cites.
Vulnerability-aware poisoning mechanism for online rl with unknown dynamics
Sun, Y. and Huang, F · 2020
Later among the works it cites.