Fetching the paper…
Reading the bibliography…
Despite the significant interest and progress in reinforcement learning (RL) problems with adversarial corruption, current works are either confined to the linear setting or lead to an undesired $\tilde{O}(\sqrt{T}\zeta)$ regret bound, where $T$ is the number of rounds and $\zeta$ is the total amount of corruption.
The epoch-greedy algorithm for multi-armed bandits with side information
Langford, J. and Zhang, T · 2007
Earlier work this paper cites.
The online loop-free stochastic shortest-path problem
Neu, G., György, A., Szepesvári, C., et al · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Linear bandits in high dimension and recommendation systems
Deshpande, Y. and Montanari, A · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D. and Van Roy, B · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Osband, I. and Van Roy, B · 2014
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Earlier work this paper cites.
Robust physical-world attacks on deep learning visual classification
Eykholt, K., Evtimov, I., Fernandes, E., Li, B., Rahmati, A., Xiao, C., Prakash, A., Kohno, T., and Song, D · 2018
Earlier work this paper cites.
Stochastic bandits robust to adversarial corruptions
Lykouris, T., Mirrokni, V., and Paes Leme, R · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Linear contextual bandits with adversarial corruptions
Zhao, H., Zhou, D., and Gu, Q · 2018
Earlier work this paper cites.
Better algorithms for stochastic bandits with adversarial corruptions
Gupta, A., Koren, T., and Talwar, K · 2019
Cited alongside, same era.
Stochastic linear optimization with adversarial corruption
Li, Y., Lou, E. Y., and Shan, L · 2019
Cited alongside, same era.
Online convex optimization in adversarial markov decision processes
Rosenberg, A. and Mansour, Y · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Wainwright, M. J · 2019
Cited alongside, same era.
Beyond ucb: Optimal and efficient contextual bandits with regression oracles
Foster, D. and Rakhlin, A · 2020
Cited alongside, same era.
Adapting to misspecification in contextual bandits
Foster, D. J., Gentile, C., Mohri, M., and Zimmert, J · 2020
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C., Liu, Q., and Miryoosefi, S · 2021
Later among the works it cites.
Online sub-sampling for reinforcement learning with general function approximation
Kong, D., Salakhutdinov, R., Wang, R., and Yang, L. F · 2021
Later among the works it cites.
Achieving near instance-optimality and minimax-optimality in stochastic and adversarial linear bandits simultaneously
Lee, C.-W., Luo, H., Wei, C.-Y., Zhang, M., and Zhang, X · 2021
Later among the works it cites.
Eluder dimension and generalized rank
Li, G., Kamath, P., Foster, D. J., and Srebro, N · 2021
Later among the works it cites.
Policy optimization in adversarial mdps: Improved exploration via dilated bonuses
Luo, H., Wei, C.-Y., and Lee, C.-W · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Simultaneously learning stochastic and adversarial episodic mdps with known transition
Jin, T. and Luo, H · 2020
Cited alongside, same era.
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Cited alongside, same era.
Stochastic shortest path with adversarially changing costs
Rosenberg, A. and Mansour, Y · 2020
Cited alongside, same era.
Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Wang, R., Salakhutdinov, R. R., and Yang, L · 2020
Cited alongside, same era.
Stochastic linear bandits robust to adversarial attacks
Bogunovic, I., Losalka, A., Krause, A., and Scarlett, J · 2021
Cited alongside, same era.
Finding the stochastic shortest path with low regret: The adversarial cost and unknown transition case
Chen, L. and Luo, H · 2021
Cited alongside, same era.
On reinforcement learning with adversarial corruption and its application to block mdp
Wu, T., Yang, Y., Du, S., and Wang, L · 2021
Later among the works it cites.
Robust stochastic linear contextual bandits under adversarial attacks
Ding, Q., Hsieh, C.-J., and Sharpnack, J · 2022
Closest in time.
Achieving minimax rates in pool-based batch active learning
Gentile, C., Wang, Z., and Zhang, T · 2022
Closest in time.
Nearly optimal algorithms for linear contextual bandits with adversarial corruptions
He, J., Zhou, D., Zhang, T., and Gu, Q · 2022
Closest in time.
A model selection approach for corruption robust reinforcement learning
Wei, C.-Y., Dann, C., and Zimmert, J · 2022
Closest in time.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Zhang, T · 2022
Closest in time.
Mathematical Analysis of Machine Learning Algorithms
Zhang, T · 2023
Closest in time.