Fetching the paper…
Reading the bibliography…
Reward learning algorithms utilize human feedback to infer a reward function, which is then used to train an AI system.
Rank analysis of incomplete block designs: I. The method of paired comparisons,
R. A. Bradley, M. E. Terry, · 1952
Earlier work this paper cites.
Algorithms for inverse reinforcement learning,
A. Y. Ng, S. J. Russell, · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning,
P. Abbeel, A. Y. Ng, · 2004
Earlier work this paper cites.
Bayesian Inverse Reinforcement Learning.,
D. Ramachandran, E. Amir, · 2007
Earlier work this paper cites.
B. D. Ziebart, Modeling purposeful adaptive behavior with the principle of maximum causal entropy, Ph.D. thesis, Carnegie Mellon University, 2010
2010
Earlier work this paper cites.
The netflix recommender system: Algorithms, business value, and innovation,
C. A. Gomez-Uribe, N. Hunt, · 2015
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search,
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al., · 2016
Earlier work this paper cites.
Learning the preferences of ignorant, inconsistent agents,
O. Evans, A. Stuhlmüller, N. D. Goodman, · 2016
Earlier work this paper cites.
Mastering the game of Go without human knowledge,
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al., · 2017
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, G. E. Hinton, · 2017
Earlier work this paper cites.
Residual attention network for image classification,
F. Wang, M. Jiang, C. Qian, S. Yang, C. Li, H. Zhang, X. Wang, X. Tang, · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions,
D. Sadigh, A. Dragan, S. Sastry, S. Seshia, · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences,
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, D. Amodei, · 2017
Earlier work this paper cites.
Learning robot objectives from physical human interaction,
A. Bajcsy, D. P. Losey, M. K. O’Malley, A. D. Dragan, · 2017
Earlier work this paper cites.
Inverse reward design,
D. Hadfield-Menell, S. Milli, P. Abbeel, S. J. Russell, A. Dragan, · 2017
Cited alongside, same era.
V. Krakovna, Specification gaming examples in AI, 2018
2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: A research direction,
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, S. Legg, · 2018
Cited alongside, same era.
S. Mindermann, R. Shah, A. Gleave, D. Hadfield-Menell, · 2018
Cited alongside, same era.
Batch active preference-based learning of reward functions,
E. Bıyık, D. Sadigh, · 2018
Cited alongside, same era.
Literal or pedagogic human? Analyzing human model misspecification in objective learning,
S. Milli, A. D. Dragan, · 2020
Later among the works it cites.
Asking easy questions: A user-friendly approach to active reward learning,
E. Bıyık, M. Palan, N. C. Landolfi, D. P. Losey, D. Sadigh, · 2020
Later among the works it cites.
Adapting a kidney exchange algorithm to align with human values,
R. Freedman, J. S. Borg, W. Sinnott-Armstrong, J. P. Dickerson, V. Conitzer, · 2020
Later among the works it cites.
A systematic study on the recommender systems in the E-commerce,
P. M. Alamdari, N. J. Navimipour, M. Hosseinzadeh, A. A. Safaei, A. Darwesh, · 2020
Later among the works it cites.
PEBBLE: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training,
K. Lee, L. M. Smith, P. Abbeel, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al., · 2019
Cited alongside, same era.
Using natural language for reward shaping in reinforcement learning,
P. Goyal, S. Niekum, R. J. Mooney, · 2019
Cited alongside, same era.
Deep reinforcement learning from policy-dependent human feedback,
D. Arumugam, J. K. Lee, S. Saskin, M. L. Littman, · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences,
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, G. Irving, · 2019
Cited alongside, same era.
Learning reward functions by integrating human demonstrations and preferences,
M. Palan, G. Shevchuk, N. Charles Landolfi, D. Sadigh, · 2019
Cited alongside, same era.
On the feasibility of learning, rather than assuming, human biases for reward inference,
R. Shah, N. Gundotra, P. Abbeel, A. Dragan, · 2019
Cited alongside, same era.
Reward-rational (implicit) choice: A unifying formalism for reward learning,
H. J. Jeon, S. Milli, A. D. Dragan, · 2020
Cited alongside, same era.
R. Freedman, R. Shah, A. Dragan, · 2021
Later among the works it cites.
Professional reviews as service: A mix method approach to assess the value of recommender systems in the entertainment industry,
M. Perano, G. L. Casali, Y. Liu, T. Abbate, · 2021
Later among the works it cites.
News recommender system: A review of recent progress, challenges, and opportunities,
S. Raza, C. Ding, · 2021
Later among the works it cites.
Human irrationality: Both bad and good for reward inference,
L. Chan, A. Critch, A. Dragan, · 2021
Later among the works it cites.
J. Leike, J. Schulman, J. Wu, Our approach to alignment research, 2022. URL: https://openai.com/blog/our-approach-to-alignment-research/
2022
Later among the works it cites.
Misspecification in inverse reinforcement learning,
J. Skalse, A. Abate, · 2022
Later among the works it cites.
The expertise problem: Learning from specialized feedback,
O. Daniels-Koch, R. Freedman, · 2022
Later among the works it cites.
Reward uncertainty for exploration in preference-based reinforcement learning,
X. Liang, K. Shu, K. Lee, P. Abbeel, · 2022
Later among the works it cites.