Fetching the paper…
Reading the bibliography…
Learning a reward function from human preferences is challenging as it typically requires having a high-fidelity simulator or using expensive and potentially unsafe actual physical rollouts in the environment.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
Offline learning from demonstrations and unlabeled experience
Zolna, K., Novikov, A., Konyushkova, K., Gulcehre, C., Wang, Z., Aytar, Y., Denil, M., de Freitas, N., and Reed, S · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Relative measurement and its generalization in decision making why pairwise comparisons are central in mathematics for the measurement of intangible factors the analytic hierarchy/network process
Saaty, T. L · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M · 2011
Earlier work this paper cites.
Batch, off-policy and model-free apprenticeship learning
Klein, E., Geist, M., and Pietquin, O · 2011
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
Zhang, J. and Cho, K · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al · 2017
Earlier work this paper cites.
Batch active preference-based learning of reward functions
Biyik, E. and Sadigh, D · 2018
Cited alongside, same era.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
Osa, T., Pajarinen, J., Neumann, G., Bagnell, J. A., Abbeel, P., and Peters, J · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, X. B., Andrychowicz, M., Zaremba, W., and Abbeel, P · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Peng, X. B., Kumar, A., Zhang, G., and Levine, S · 2019
Later among the works it cites.
Derail: Diagnostic environments for reward and imitation learning
Freire, P., Gleave, A., Toyer, S., and Russell, S · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Jeon, H. J., Milli, S., and Dragan, A. D · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Brown, D., Goo, W., Nagarajan, P., and Niekum, S · 2019
Cited alongside, same era.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Cabi, S., Colmenarejo, S. G., Novikov, A., Konyushkova, K., Reed, S., Jeong, R., Zolna, K., Aytar, Y., Budden, D., Vecerik, M., et al · 2019
Cited alongside, same era.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Chebotar, Y., Handa, A., Makoviychuk, V., Macklin, M., Issac, J., Ratliff, N., and Fox, D · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
Risk-sensitive generative adversarial imitation learning
Lacotte, J., Ghavamzadeh, M., Chow, Y., and Pavone, M · 2019
Cited alongside, same era.
Truly batch apprenticeship learning with deep successor features
Lee, D., Srinivasan, S., and Doshi-Velez, F · 2019
Cited alongside, same era.
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Later among the works it cites.
Learning human objectives by evaluating hypothetical behavior
Reddy, S., Dragan, A., Levine, S., Legg, S., and Leike, J · 2020
Later among the works it cites.
The magical benchmark for robust imitation
Toyer, S., Shah, R., Critch, A., and Russell, S · 2020
Later among the works it cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
Arora, S. and Doshi, P · 2021
Closest in time.
Learning” what-if” explanations for sequential decision-making
Bica, I., Jarrett, D., Hüyük, A., and van der Schaar, M · 2021
Closest in time.
Replacing rewards with examples: Example-based policy search via recursive classification
Eysenbach, B., Levine, S., and Salakhutdinov, R · 2021
Closest in time.
Thriftydagger: Budget-aware novelty and risk gating for interactive imitation learning
Hoque, R., Balakrishna, A., Novoseller, E., Wilcox, A., Brown, D. S., and Goldberg, K · 2021
Closest in time.
Lee, K., Smith, L., and Abbeel, P · 2021
Closest in time.
Self-supervised online reward shaping in sparse-reward environments
Memarian, F., Goo, W., Lioutikov, R., Topcu, U., and Niekum, S · 2021
Closest in time.