Fetching the paper…
Reading the bibliography…
A core element in decision-making under uncertainty is the feedback on the quality of the performed actions.
Estimation des densités: risque minimax
Bretagnolle, J. and Huber, C · 1979
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
Dani, V., Hayes, T. P., and Kakade, S. M · 2008
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M · 2009
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C · 2011
Earlier work this paper cites.
Contextual bandit algorithms with supervised learning guarantees
Beygelzimer, A., Langford, J., Li, L., Reyzin, L., and Schapire, R · 2011
Earlier work this paper cites.
The kl-ucb algorithm for bounded stochastic bandits and beyond
Garivier, A. and Cappé, O · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Earlier work this paper cites.
Thompson sampling: An asymptotically optimal finite-time analysis
Kaufmann, E., Korda, N., and Munos, R · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N · 2013
Earlier work this paper cites.
Bandits with knapsacks
Badanidiyuru, A., Kleinberg, R., and Slivkins, A · 2013
Earlier work this paper cites.
How hard is my mdp?” the distribution-norm to the rescue”
Maillard, O.-A., Mann, T. A., and Mannor, S · 2014
Earlier work this paper cites.
Prediction with limited advice and multiarmed bandits with paid observations
Seldin, Y., Bartlett, P., Crammer, K., and Abbasi-Yadkori, Y · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Anytime optimal algorithms in stochastic multi-armed bandits
Degenne, R. and Perchet, V · 2016
Cited alongside, same era.
Linear thompson sampling revisited
Abeille, M., Lazaric, A., et al · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
Learning adversarial mdps with bandit feedback and unknown transition
Jin, T. and Luo, H · 2019
Later among the works it cites.
Batch-size independent regret bounds for the combinatorial multi-armed bandit problem
Merlis, N. and Mannor, S · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. G · 2019
Later among the works it cites.
Zanette, A. and Brunskill, E · 2019
Later among the works it cites.
Near-optimal regret bounds for stochastic shortest path
Cohen, A., Kaplan, H., Mansour, Y., and Rosenberg, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Multi-armed bandit with additional observations
Yun, D., Proutiere, A., Ahn, S., Shin, J., and Yi, Y · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E · 2019
Cited alongside, same era.
Tight regret bounds for model-based reinforcement learning with greedy policies
Efroni, Y., Merlis, N., Ghavamzadeh, M., and Mannor, S · 2019
Cited alongside, same era.
Model selection for contextual bandits
Foster, D. J., Krishnamurthy, A., and Luo, H · 2019
Cited alongside, same era.
Explore first, exploit next: The true shape of regret in bandit problems
Garivier, A., Ménard, P., and Stoltz, G · 2019
Cited alongside, same era.
Reinforcement learning with trajectory feedback
Efroni, Y., Merlis, N., and Mannor, S · 2020
Later among the works it cites.
Foster, D. J., Rakhlin, A., Simchi-Levi, D., and Xu, Y · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Later among the works it cites.
Tight lower bounds for combinatorial multi-armed bandits
Merlis, N. and Mannor, S · 2020
Later among the works it cites.
No-regret exploration in goal-oriented reinforcement learning
Tarbouriech, J., Garcelon, E., Valko, M., Pirotta, M., and Lazaric, A · 2020
Later among the works it cites.
Zhang, Z., Ji, X., and Du, S. S · 2020
Later among the works it cites.
Multi-armed bandits with cost subsidy
Sinha, D., Sankararaman, K. A., Kazerouni, A., and Avadhanula, V · 2021
Closest in time.