Fetching the paper…
Reading the bibliography…
We study human-in-the-loop reinforcement learning (RL) with trajectory preferences, where instead of receiving a numeric reward at each step, the agent only receives preferences over trajectory pairs from a human overseer.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Wang, R., Salakhutdinov, R., and Yang, L. F · 2005
Earlier work this paper cites.
On reward-free reinforcement learning with linear function approximation
Wang, R., Du, S. S., Yang, L. F., and Salakhutdinov, R · 2006
Earlier work this paper cites.
Preference-based reinforcement learning with finite-time guarantees
Xu, Y., Wang, R., Yang, L. F., Singh, A., and Dubrawski, A · 2006
Earlier work this paper cites.
Reinforcement learning strategies for clinical trials in nonsmall cell lung cancer
Zhao, Y., Zeng, D., Socinski, M. A., and Kosorok, M. R · 2011
Earlier work this paper cites.
First steps towards learning from game annotations
Wirth, C. and Fürnkranz, J · 2012
Earlier work this paper cites.
The k-armed dueling bandits problem
Yue, Y., Broder, J., Kleinberg, R., and Joachims, T · 2012
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Jain, A., Wojcik, B., Joachims, T., and Saxena, A · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Russo, D. and Van Roy, B · 2013
Earlier work this paper cites.
Preference-based reinforcement learning: evolutionary direct policy search using a preference-based racing algorithm
Busa-Fekete, R., Szörényi, B., Weng, P., Cheng, W., and Hüllermeier, E · 2014
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
Osband, I. and Van Roy, B · 2014
Earlier work this paper cites.
Learning to optimize via posterior sampling
Russo, D. and Van Roy, B · 2014
Earlier work this paper cites.
On learning from game annotations
Wirth, C. and Fürnkranz, J · 2014
Cited alongside, same era.
Learning preferences for manipulation tasks from online coactive feedback
Jain, A., Sharma, S., Joachims, T., and Saxena, A · 2015
Cited alongside, same era.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Sample-optimal parametric q-learning using linearly additive features
Yang, L. and Wang, M · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A., Jia, Z., Szepesvari, C., Wang, M., and Yang, L · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z · 2020
Later among the works it cites.
Reinforcement learning with trajectory feedback
Efroni, Y., Merlis, N., and Mannor, S · 2020
Later among the works it cites.
Foster, D. J., Rakhlin, A., Simchi-Levi, D., and Xu, Y · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Cited alongside, same era.
Imitation learning: A survey of learning methods
Hussein, A., Gaber, M. M., Elyan, E., and Jayne, C · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al · 2017
Cited alongside, same era.
Preference-based online learning with dueling bandits: A survey
Busa-Fekete, R., Hüllermeier, E., and Mesaoudi-Paul, A. E · 2018
Cited alongside, same era.
Learning from comparisons and choices
Negahban, S., Oh, S., Thekumparampil, K. K., and Xu, J · 2018
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I · 2020
Later among the works it cites.
Dueling posterior sampling for preference-based reinforcement learning
Novoseller, E., Wei, Y., Sui, Y., Yue, Y., and Burdick, J · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
Zanette, A., Lazaric, A., Kochenderfer, M., and Brunskill, E · 2020
Later among the works it cites.
Bayesian optimization with safety constraints: safe and automatic parameter tuning in robotics
Berkenkamp, F., Krause, A., and Schoellig, A. P · 2021
Later among the works it cites.
On the theory of reinforcement learning with once-per-episode feedback
Chatterji, N. S., Pacchiano, A., Bartlett, P. L., and Jordan, M. I · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S., Kakade, S. M., Lee, J. D., Lovett, S., Mahajan, G., Sun, W., and Wang, R · 2021
Later among the works it cites.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Jin, C., Liu, Q., and Miryoosefi, S · 2021
Later among the works it cites.
Online sub-sampling for reinforcement learning with general function approximation
Kong, D., Salakhutdinov, R., Wang, R., and Yang, L. F · 2021
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences
Pacchiano, A., Saha, A., and Lee, J · 2021
Later among the works it cites.
Reward-free model-based reinforcement learning with linear function approximation
Zhang, W., Zhou, D., and Gu, Q · 2021
Later among the works it cites.