Fetching the paper…
Reading the bibliography…
We investigate the extent to which offline demonstration data can improve online learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Multi-armed bandit problems with history
Shivaswamy, P. and Joachims, T · 2012
Earlier work this paper cites.
Prior-free and prior-dependent regret bounds for thompson sampling
Bubeck, S. and Liu, C.-Y · 2013
Earlier work this paper cites.
Learning to optimize via information-directed sampling
Russo, D. and Van Roy, B · 2014
Earlier work this paper cites.
CVXPY: A Python-embedded modeling language for convex optimization
Diamond, S. and Boyd, S · 2016
Earlier work this paper cites.
An information-theoretic analysis of thompson sampling
Russo, D. and Van Roy, B · 2016
Earlier work this paper cites.
Ensemble sampling
Lu, X. and Van Roy, B · 2017
Earlier work this paper cites.
Thompson sampling for stochastic bandits with graph feedback
Tossou, A. C., Dimitrakakis, C., and Dubhashi, D · 2017
Earlier work this paper cites.
Bandits with side observations: Bounded vs. logarithmic regret
Degenne, R., Garcelon, E., and Perchet, V · 2018
Earlier work this paper cites.
Randomized prior functions for deep reinforcement learning
Osband, I., Aslanides, J., and Cassirer, A · 2018
Earlier work this paper cites.
A tutorial on thompson sampling
Russo, D. J., Van Roy, B., Kazerouni, A., Osband, I., Wen, Z., et al · 2018
Cited alongside, same era.
An information-theoretic approach to minimax regret in partial monitoring
Lattimore, T. and Szepesvári, C · 2019
Cited alongside, same era.
Deep exploration via randomized value functions
Osband, I., Van Roy, B., Russo, D. J., Wen, Z., et al · 2019
Cited alongside, same era.
Warm-starting contextual bandits: Robustly combining supervised and bandit feedback
Zhang, C., Agarwal, A., Daumé III, H., Langford, J., and Negahban, S. N · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
Artificial replay: A meta-algorithm for harnessing historical data in bandits
Banerjee, S., Sinclair, S. R., Tambe, M., Xu, L., and Yu, C. L · 2022
Later among the works it cites.
Imitation learning by estimating expertise of demonstrators
Beliaev, M., Shih, A., Ermon, S., Sadigh, D., and Pedarsani, R · 2022
Later among the works it cites.
Leveraging initial hints for free in stochastic linear bandits
Cutkosky, A., Dann, C., Das, A., and Zhang, Q · 2022
Later among the works it cites.
Ensembles for uncertainty estimation: Benefits of prior functions and bootstrapping
Dwaracherla, V., Wen, Z., Osband, I., Lu, X., Asghari, S. M., and Van Roy, B · 2022
Later among the works it cites.
Regret bounds for information-directed reinforcement learning
Hao, B. and Lattimore, T · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bandit algorithms
Lattimore, T. and Szepesvári, C · 2020
Cited alongside, same era.
Information directed sampling for sparse linear bandits
Hao, B., Lattimore, T., and Deng, W · 2021
Cited alongside, same era.
Meta-thompson sampling
Kveton, B., Konobeev, M., Zaheer, M., Hsu, C.-w., Mladenov, M., Boutilier, C., and Szepesvari, C · 2021
Cited alongside, same era.
Osband, I., Wen, Z., Asghari, M., Ibrahimi, M., Lu, X., and Van Roy, B · 2021
Cited alongside, same era.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S · 2021
Cited alongside, same era.
Bayesian decision-making under misspecified priors with applications to meta-learning
Simchowitz, M., Tosh, C., Krishnamurthy, A., Hsu, D. J., Lykouris, T., Dudik, M., and Schapire, R. E · 2021
Cited alongside, same era.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y · 2021
Cited alongside, same era.
Osband, I., Wen, Z., Asghari, S. M., Dwaracherla, V., Lu, X., Ibrahimi, M., Lawson, D., Hao, B., O’Donoghue, B., and Roy, B. V · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
An analysis of ensemble sampling
Qin, C., Wen, Z., Lu, X., and Van Roy, B · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W · 2022
Later among the works it cites.
Leveraging offline data in online reinforcement learning
Wagenmaker, A. and Pacchiano, A · 2022
Later among the works it cites.
Feel-good thompson sampling for contextual bandits and reinforcement learning
Zhang, T · 2022
Later among the works it cites.