Fetching the paper…
Reading the bibliography…
Reward functions are notoriously difficult to specify, especially for tasks with complex goals.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. F. Christiano, and G. Irving · 1909
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
R. A. Bradley and M. E. Terry · 1952
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
D. Ramachandran and E. Amir · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
B. D. Ziebart, J. A. Bagnell, and A. K. Dey · 2010
Earlier work this paper cites.
Preference-based policy learning
R. Akrour, M. Schoenauer, and M. Sebag · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
A Bayesian approach for policy learning from trajectory preference queries
A. Wilson, A. Fern, and P. Tadepalli · 2012
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Cited alongside, same era.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. Sastry, and S. A. Seshia · 2017
Cited alongside, same era.
Learning robust rewards with adversarial inverse reinforcement learning
J. Fu, K. Luo, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine · 2018
F-IRL: Inverse Reinforcement Learning via State Marginal Matching, Dec. 2020
T. Ni, H. Sikchi, Y. Wang, T. Gupta, L. Lee, and B. Eysenbach · 2020
Later among the works it cites.
Learning human objectives by evaluating hypothetical behavior, 2020
S. Reddy, A. D. Dragan, S. Levine, S. Legg, and J. Leike · 2020
Later among the works it cites.
The imitation
S. Wang, S. Toyer, A. Gleave, and S. Emmons · 2020
Later among the works it cites.
Consequences of misaligned AI
S. Zhuang and D. Hadfield-Menell · 2020
Later among the works it cites.
Quantifying differences in reward functions
A. Gleave, M. D. Dennis, S. Legg, S. Russell, and J. Leike · 2021
Later among the works it cites.
PEBBLE: feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
K. Lee, L. M. Smith, and P. Abbeel · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reward learning from human preferences and demonstrations in Atari
B. Ibarz, J. Leike, T. Pohlen, G. Irving, S. Legg, and D. Amodei · 2018
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. S. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Cited alongside, same era.
Random expert distillation: Imitation learning via expert policy support estimation
R. Wang, C. Ciliberto, P. V. Amadori, and Y. Demiris · 2019
Cited alongside, same era.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
S. Cabi, S. Gómez Colmenarejo, A. Novikov, K. Konyushova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, O. Sushkov, D. Barker, J. Scholz, M. Denil, N. de Freitas, and Z. Wang · 2020
Cited alongside, same era.
seals: Suite of environments for algorithms that learn specifications
A. Gleave, P. Freire, S. Wang, and S. Toyer · 2020
Cited alongside, same era.
Reward-rational (implicit) choice: A unifying formalism for reward learning
H. J. Jeon, S. Milli, and A. D. Dragan · 2020
Cited alongside, same era.
B-pref: Benchmarking preference-based reinforcement learning
K. Lee, L. M. Smith, A. D. Dragan, and P. Abbeel · 2021
Later among the works it cites.
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Later among the works it cites.
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback, Apr. 2022
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan · 2022
Later among the works it cites.
WebGPT: Browser-assisted question-answering with human feedback, June 2022
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback, Mar. 2022
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
A. Pan, K. Bhatia, and J. Steinhardt · 2022
Later among the works it cites.
Learning to summarize from human feedback, Feb. 2022
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano · 2022
Later among the works it cites.