Fetching the paper…
Reading the bibliography…
In this paper, we investigate the problem of offline Preference-based Reinforcement Learning (PbRL) with human feedback where feedback is available in the form of preference between trajectory pairs rather than explicit rewards.
Empirical Processes in M-estimation
van de Geer, S. (2000) · 2000
Earlier work this paper cites.
Fast learning rates for plug-in classifiers
Audibert, J.-Y. and Tsybakov, A. B. (2007) · 2007
Earlier work this paper cites.
Provably good batch reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E. (2020) · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Ramachandran, D. and Amir, E. (2007) · 2007
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z. (2020) · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T. (2012) · 2012
Earlier work this paper cites.
The multi-armed bandit problem with covariates
Perchet, V. and Rigollet, P. (2013) · 2013
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L. (2017) · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al. (2017) · 2017
Earlier work this paper cites.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Warnell, G., Waytowich, N., Lawhern, V., and Stone, P. (2018) · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Agarwal, A., Jiang, N., Kakade, S. M., and Sun, W. (2019) · 2019
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D., Goo, W., Nagarajan, P., and Niekum, S. (2019) · 2019
Earlier work this paper cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. G. (2019) · 2019
Earlier work this paper cites.
High-Dimensional Statistics: A Non-Asymptotic Viewpoint
Wainwright, M. (2019) · 2019
Earlier work this paper cites.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T. (2020) · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S. (2020) · 2020
Cited alongside, same era.
Performance guarantees for policy learning
Luedtke, A. and Chambaz, A. (2020) · 2020
Cited alongside, same era.
Dueling posterior sampling for preference-based reinforcement learning
Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J. (2020) · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020) · 2020
Cited alongside, same era.
Preference-based reinforcement learning with finite-time guarantees
Xu, Y., Wang, R., Yang, L., Singh, A., and Dubrawski, A. (2020) · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T. (2020) · 2020
Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation
Chen, X., Zhong, H., Yang, Z., Wang, Z., and Wang, L. (2022) · 2022
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Cheng, C.-A., Xie, T., Jiang, N., and Agarwal, A. (2022) · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trębacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. (2022) · 2022
Later among the works it cites.
Fast rates for contextual linear optimization
Hu, Y., Kallus, N., and Mao, X. (2022) · 2022
Later among the works it cites.
When is partially observable reinforcement learning not scary?
Liu, Q., Chung, A., Szepesvári, C., and Jin, C. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast rates for the regret of offline reinforcement learning
Hu, Y., Kallus, N., and Uehara, M. (2021) · 2021
Cited alongside, same era.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. (2021) · 2021
Cited alongside, same era.
Dueling rl: reinforcement learning with trajectory preferences
Pacchiano, A., Saha, A., and Lee, J. (2021) · 2021
Cited alongside, same era.
Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage
Uehara, M. and Sun, W. (2021) · 2021
Cited alongside, same era.
Recursively summarizing books with human feedback
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P. (2021) · 2021
Cited alongside, same era.
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y. (2022) · 2022
Later among the works it cites.
Rambo-rl: Robust adversarial model-based offline reinforcement learning
Rigter, M., Lacerda, B., and Hawes, N. (2022) · 2022
Later among the works it cites.
Pessimistic q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Later among the works it cites.
Hybrid rl: Using both offline and online data can make rl efficient
Song, Y., Zhou, Y., Sekhari, A., Bagnell, J. A., Krishnamurthy, A., and Sun, W. (2022) · 2022
Later among the works it cites.
Gap-dependent unsupervised exploration for reinforcement learning
Wu, J., Braverman, V., and Yang, L. (2022) · 2022
Later among the works it cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2019) · 2022
Later among the works it cites.
Languages are rewards: Hindsight finetuning using human feedback
Liu, H., Sferrazza, C., and Abbeel, P. (2023) · 2023
Closest in time.
Benchmarks and algorithms for offline preference-based reward learning
Shin, D., Dragan, A. D., and Brown, D. S. (2023) · 2023
Closest in time.
Refined value-based offline rl under realizability and partial coverage
Uehara, M., Kallus, N., Lee, J. D., and Sun, W. (2023) · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Zhu, B., Jiao, J., and Jordan, M. I. (2023) · 2023
Closest in time.