Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) is a variant of reinforcement learning (RL) that learns from human feedback instead of relying on an engineered reward function.
Deep Reinforcement Learning from Policy-Dependent Human Feedback, 2019
Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L. Littman · 1902
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences, 2020
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 1909
Earlier work this paper cites.
A law of comparative judgment
Louis Leon Thurstone · 1927
Earlier work this paper cites.
Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
Individual Choice Behavior: A Theoretical Analysis
R. Duncan Luce · 1959
Earlier work this paper cites.
A Theory of the Learnable
L. G. Valiant · 1972
Earlier work this paper cites.
The Analysis of Permutations
R. L. Plackett · 1975
Earlier work this paper cites.
Optimal Replacement of GMC Bus Engines: An Empirical Model of Harold Zurcher
John Rust · 1987
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y. Ng and Stuart J. Russell · 2000
Earlier work this paper cites.
Eligibility Traces for Off-Policy Policy Evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
A social reinforcement learning agent
Charles Isbell, Christian R. Shelton, Michael Kearns, Satinder Singh, and Peter Stone · 2001
Earlier work this paper cites.
Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression
Richard Maclin, Jude Shavlik, Lisa Torrey, Trevor Walker, and Edward Wild · 2005
Earlier work this paper cites.
The Construction of Preference
Sarah Lichtenstein and Paul Slovic (eds.) · 2006
Earlier work this paper cites.
Risk bounds for statistical learning
Pascal Massart and Élodie Nédélec · 2006
Earlier work this paper cites.
Differential Privacy: A Survey of Results
Cynthia Dwork · 2008
Earlier work this paper cites.
Using Discrete Choice Experiments to Value Health and Health Care , volume 11
Mandy Ryan, Karen Gerard, Mabel Amaya-Amaya, and Ian J. Bateman (eds.) · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W. Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Training parsers by inverse reinforcement learning
Gergely Neu and Csaba Szepesvári · 2009
Earlier work this paper cites.
Regret-based reward elicitation for Markov decision processes
Kevin Regan and Craig Boutilier · 2009
Earlier work this paper cites.
Discrete Choice Methods with Simulation
Kenneth E. Train · 2009
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yisong Yue and Thorsten Joachims · 2009
Earlier work this paper cites.
Inverse optimal control with linearly-solvable MDPs
Krishnamurthy Dvijotham and Emanuel Todorov · 2010
Earlier work this paper cites.
Reinforcement Learning Via Practice and Critique Advice
Kshitij Judah, Saikat Roy, Alan Fern, and Thomas Dietterich · 2010
Earlier work this paper cites.
Preference-Based Policy Learning
Riad Akrour, Marc Schoenauer, and Michele Sebag · 2011
Earlier work this paper cites.
Preference-Based Policy Iteration: Leveraging Preference Learning for Reinforcement Learning
Weiwei Cheng, Johannes Fürnkranz, Eyke Hüllermeier, and Sang-Hyeun Park · 2011
Earlier work this paper cites.
Ordering effects and choice set awareness in repeat-response stated preference studies
Brett Day, Ian J. Bateman, Richard T. Carson, Diane Dupont, Jordan J. Louviere, Sanae Morimoto, Riccardo Scarpa, and Paul Wang · 2011
Earlier work this paper cites.
Gibbs Measures and Phase Transitions
Hans-Otto Georgii · 2011
Earlier work this paper cites.
Portfolio Allocation for Bayesian Optimization
Matthew Hoffman, Eric Brochu, and Nando de Freitas · 2011
Earlier work this paper cites.
Consumer decision making in knowledge-based recommendation
Monika Mandl, Alexander Felfernig, Erich Teppan, and Monika Schubert · 2011
Earlier work this paper cites.
Robust online optimization of reward-uncertain MDPs
Kevin Regan and Craig Boutilier · 2011
Earlier work this paper cites.
Models for Paired Comparison Data: A Review with Emphasis on Dependent Data
Manuela Cattelan · 2012
Earlier work this paper cites.
Nonparametric Bayesian Inverse Reinforcement Learning for Multiple Reward Functions
Jaedeug Choi and Kee-eung Kim · 2012
Earlier work this paper cites.
Preference-based reinforcement learning: A formal framework and a policy iteration algorithm
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng, and Sang-Hyeun Park · 2012
Earlier work this paper cites.
Learning from Human-generated Reward
W. Bradley Knox · 2012
Earlier work this paper cites.
Designing interfaces for explicit preference elicitation: A user-centered investigation of preference representation and elicitation process
Alina Pommeranz, Joost Broekens, Pascal Wiggers, Willem-Paul Brinkman, and Catholijn M. Jonker · 2012
Earlier work this paper cites.
Active Learning
Burr Settles · 2012
Earlier work this paper cites.
Online structured prediction via coactive learning
Pannaga Shivaswamy and Thorsten Joachims · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
Hiroaki Sugiyama, Toyomi Meguro, and Yasuhiro Minami · 2012
Earlier work this paper cites.
A Bayesian Approach for Policy Learning from Trajectory Preference Queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Human Decision Making and Recommender Systems
Li Chen, Marco de Gemmis, Alexander Felfernig, Pasquale Lops, Francesco Ricci, and Giovanni Semeraro · 2013
Earlier work this paper cites.
Survey Research Methods
Floyd J. Fowler, Jr · 2013
Earlier work this paper cites.
Off-Policy Evaluation in Markov Decision Processes
Cosmin Paduraru · 2013
Earlier work this paper cites.
Eluder Dimension and the Sample Complexity of Optimistic Exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Interactive value iteration for Markov decision processes with unknown rewards
Paul Weng and Bruno Zanuttini · 2013
Earlier work this paper cites.
Statistical Methods for Ranking Data
Mayer Alvo and Philip L.H. Yu · 2014
Earlier work this paper cites.
Preference-based reinforcement learning: Evolutionary direct policy search using a preference-based racing algorithm
Róbert Busa-Fekete, Balázs Szörényi, Paul Weng, Weiwei Cheng, and Eyke Hüllermeier · 2014
Earlier work this paper cites.
Active Reward Learning
Christian Daniel, Malte Viering, Jan Metz, Oliver Kroemer, and Jan Peters · 2014
Earlier work this paper cites.
Random Choice as Behavioral Optimization
Faruk Gul, Paulo Natenzon, and Wolfgang Pesendorfer · 2014
Earlier work this paper cites.
Programming by Feedback
Marc Schoenauer, Riad Akrour, Michele Sebag, and Jean-Christophe Souplet · 2014
Earlier work this paper cites.
Contextual Dueling Bandits
Miroslav Dudík, Katja Hofmann, Robert E. Schapire, Aleksandrs Slivkins, and Masrour Zoghi · 2015
Earlier work this paper cites.
Reducing the Number of Queries in Interactive Value Iteration
Hugo Gilbert, Olivier Spanjaard, Paolo Viappiani, and Paul Weng · 2015
Earlier work this paper cites.
Learning preferences for manipulation tasks from online coactive feedback
Ashesh Jain, Shikhar Sharma, Thorsten Joachims, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Coactive Learning
Pannaga Shivaswamy and Thorsten Joachims · 2015
Earlier work this paper cites.
Ratings are Overrated!
Georgios N. Yannakakis and Héctor P. Martínez · 2015
Earlier work this paper cites.
Concrete Problems in AI Safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Handbook of Computational Social Choice
Felix Brandt, Vincent Conitzer, Ulle Endriss, Jérôme Lang, and Ariel D. Procaccia (eds.) · 2016
Earlier work this paper cites.
Faulty Reward Functions in the Wild, 2016
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
Quantile Reinforcement Learning
Hugo Gilbert and Paul Weng · 2016
Earlier work this paper cites.
Model-Free Reinforcement Learning with Skew-Symmetric Bilinear Utilities
Hugo Gilbert, Bruno Zanuttini, Paolo Viappiani, Paul Weng, and Esther Nicart · 2016
Earlier work this paper cites.
Cooperative Inverse Reinforcement Learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Active comparison based learning incorporating user uncertainty and noise
Rachel Holladay, Shervin Javdani, Anca Dragan, and Siddhartha Srinivasa · 2016
Earlier work this paper cites.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Model-Free Preference-Based Reinforcement Learning
Christian Wirth, Johannes Fürnkranz, and Gerhard Neumann · 2016
Earlier work this paper cites.
Double Thompson Sampling for Dueling Bandits
Huasen Wu and Xin Liu · 2016
Earlier work this paper cites.
Do You Want Your Autonomous Car To Drive Like You?
Chandrayee Basu, Qian Yang, David Hungerman, Mukesh Singhal, and Anca D. Dragan · 2017
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Optimizing Quantiles in Preference-Based Markov Decision Processes
Hugo Gilbert, Paul Weng, and Yan Xu · 2017
Earlier work this paper cites.
Reinforcement Learning with Deep Energy-Based Policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
The Off-Switch Game
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2017
Earlier work this paper cites.
Inverse Reward Design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Earlier work this paper cites.
Grounded Language Learning in a Simulated 3D World, 2017
Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojciech Marian Czarnecki, Max Jaderberg, Denis Teplyashin, Marcus Wainwright, Chris Apps, Demis Hassabis, and Phil Blunsom · 2017
Earlier work this paper cites.
Sequence Tutor: Conservative Fine-Tuning of Sequence Generation Models with KL-control
Natasha Jaques, Shixiang Gu, Dzmitry Bahdanau, José Miguel Hernández-Lobato, Richard E. Turner, and Douglas Eck · 2017
Earlier work this paper cites.
A Learning Theory of Ranking Aggregation
Anna Korba, Stéphan Clemencon, and Eric Sibony · 2017
Earlier work this paper cites.
Interactive Learning from Policy-Dependent Human Feedback
James MacGlashan, Mark K. Ho, Robert Loftin, Bei Peng, Guan Wang, David L. Roberts, Matthew E. Taylor, and Michael L. Littman · 2017
Earlier work this paper cites.
Active Preference-Based Learning of Reward Functions
Dorsa Sadigh, Anca Dragan, Shankar Sastry, and Sanjit Seshia · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Multi-dueling Bandits with Dependent Arms
Yanan Sui, Vincent Zhuang, Joel W. Burdick, and Yisong Yue · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A Survey of Preference-Based Reinforcement Learning Methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Earlier work this paper cites.
Learning from Richer Human Guidance: Augmenting Comparison-Based Learning with Feature Queries
Chandrayee Basu, Mukesh Singhal, and Anca D. Dragan · 2018
Earlier work this paper cites.
Batch Active Preference-Based Learning of Reward Functions
Erdem Bıyık and Dorsa Sadigh · 2018
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts, 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
Active Reward Learning from Critiques
Yuchen Cui and Scott Niekum · 2018
Earlier work this paper cites.
Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
Justin Fu, Avi Singh, Dibya Ghosh, Larry Yang, and Sergey Levine · 2018
Earlier work this paper cites.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
APRIL: Interactively Learning to Summarise by Combining Active Preference Learning and Reinforcement Learning
Yang Gao, Christian M. Meyer, and Iryna Gurevych · 2018
Earlier work this paper cites.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Rainbow: Combining Improvements in Deep Reinforcement Learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in Atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Interaction Algorithm Effect on Human Experience with Reinforcement Learning
Samantha Krening and Karen M. Feigh · 2018
Earlier work this paper cites.
Learning Dynamic Robot-to-Human Object Handover from Human Feedback
Andras Kupcsik, David Hsu, and Wee Sun Lee · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: A research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Including Uncertainty when Learning from Human Corrections
Dylan P. Losey and Marcia K. O’Malley · 2018
Earlier work this paper cites.
Lifelong Inverse Reinforcement Learning
Jorge Mendez, Shashank Shivkumar, and Eric Eaton · 2018
Earlier work this paper cites.
Occam’ s razor is insufficient to infer the preferences of irrational agents
Sören Mindermann and Stuart Armstrong · 2018
Earlier work this paper cites.
Active Inverse Reward Design
Sören Mindermann, Rohin Shah, Adam Gleave, and Dylan Hadfield-Menell · 2018
Earlier work this paper cites.
An Algorithmic Perspective on Imitation Learning
Takayuki Osa, Joni Pajarinen, Gerhard Neumann, J. Andrew Bagnell, Pieter Abbeel, and Jan Peters · 2018
Earlier work this paper cites.
Sample and Feedback Efficient Hierarchical Reinforcement Learning from Human Preferences
Robert Pinsler, Riad Akrour, Takayuki Osa, Jan Peters, and Gerhard Neumann · 2018
Earlier work this paper cites.
Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
William Saunders, Girish Sastry, Andreas Stuhlmüller, and Owain Evans · 2018
Earlier work this paper cites.
Advancements in Dueling Bandits
Yanan Sui, Masrour Zoghi, Katja Hofmann, and Yisong Yue · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Earlier work this paper cites.
Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2018
Earlier work this paper cites.
Learning User Preferences in Robot Motion Planning Through Interaction
Nils Wilde, Dana Kulić, and Stephen L. Smith · 2018
Earlier work this paper cites.
Few-Shot Goal Inference for Visuomotor Learning and Planning
Annie Xie, Avi Singh, Sergey Levine, and Chelsea Finn · 2018
Earlier work this paper cites.
Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera · 2019
Earlier work this paper cites.
Active Learning of Reward Dynamics from Hierarchical Queries
Chandrayee Basu, Erdem Bıyık, Zhixun He, Mukesh Singhal, and Dorsa Sadigh · 2019
Earlier work this paper cites.
Asking Easy Questions: A User-Friendly Approach to Active Reward Learning
Erdem Bıyık, Malayandi Palan, Nicholas C. Landolfi, Dylan P. Losey, and Dorsa Sadigh · 2019
Earlier work this paper cites.
Better Rewards Yield Better Summaries: Learning to Summarise Without References
Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych · 2019
Earlier work this paper cites.
Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Earlier work this paper cites.
Bridging the gap between regret minimization and best arm identification, with application to A/B tests
Rémy Degenne, Thomas Nedelec, Clement Calauzenes, and Vianney Perchet · 2019
Earlier work this paper cites.
DeepMDP: Learning Continuous Latent Space Models for Representation Learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Earlier work this paper cites.
Off-Policy Evaluation via Off-Policy Classification
Alexander Irpan, Kanishka Rao, Konstantinos Bousmalis, Chris Harris, Julian Ibarz, and Sergey Levine · 2019
Earlier work this paper cites.
Batch Policy Learning under Constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Earlier work this paper cites.
A Survey of Reinforcement Learning Informed by Natural Language
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 2019
Earlier work this paper cites.
Learning Reward Functions by Integrating Human Demonstrations and Preferences
Malayandi Palan, Gleb Shevchuk, Nicholas Charles Landolfi, and Dorsa Sadigh · 2019
Earlier work this paper cites.
Teacher-Aware Active Robot Learning
Mattia Racca, Antti Oulasvirta, and Ville Kyrki · 2019
Earlier work this paper cites.
Preferences Implicit in the State of the World
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
End-To-End Robotic Reinforcement Learning without Reward Engineering
Avi Singh, Larry Yang, Chelsea Finn, and Sergey Levine · 2019
Earlier work this paper cites.
Pitfalls of Learning a Reward Function Online
Stuart Armstrong, Jan Leike, Laurent Orseau, and Shane Legg · 2020
Earlier work this paper cites.
Emergent Tool Use From Multi-Agent Autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2020
Earlier work this paper cites.
Preselection Bandits
Viktor Bengs and Eyke Hüllermeier · 2020
Earlier work this paper cites.
Active Preference-Based Gaussian Process Regression for Reward Learning
Erdem Bıyık, Nicolas Huynh, Mykel Kochenderfer, and Dorsa Sadigh · 2020
Earlier work this paper cites.
Quantifying Hypothesis Space Misspecification in Learning From Human–Robot Demonstrations and Physical Corrections
Andreea Bobu, Andrea Bajcsy, Jaime F. Fisac, Sampada Deglurkar, and Anca D. Dragan · 2020
Earlier work this paper cites.
Safe Imitation Learning via Fast Bayesian Reward Inference from Preferences
Daniel S. Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum · 2020
Earlier work this paper cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, Oleg Sushkov, David Barker, Jonathan Scholz, Misha Denil, Nando de Freitas, and Ziyu Wang · 2020
Earlier work this paper cites.
A Survey on Interactive Reinforcement Learning: Design Principles and Open Challenges
Christian Arzate Cruz and Takeo Igarashi · 2020
Earlier work this paper cites.
DERAIL: Diagnostic Environments for Reward And Imitation Learning
Pedro Freire, Adam Gleave, Sam Toyer, and Stuart Russell · 2020
Earlier work this paper cites.
Artificial Intelligence, Values, and Alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Preference-based interactive multi-document summarisation
Yang Gao, Christian M. Meyer, and Iryna Gurevych · 2020
Cited alongside, same era.
Generalized transitivity: A systematic comparison of concepts with an application to preferences in the Babington Smith model
Björn Haddenhorst, Eyke Hüllermeier, and Martin Kolb · 2020
Cited alongside, same era.
Dream to Control: Learning Behaviors by Latent Imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2020
Cited alongside, same era.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2020
Cited alongside, same era.
Preference-Based Image Generation
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, KaShun Shum, and Tong Zhang · 2023
Closest in time.
Vision-Language Models as Success Detectors
Yuqing Du, Ksenia Konyushkova, Misha Denil, Akhil Raju, Jessica Landon, Felix Hill, Nando de Freitas, and Serkan Cabi · 2023
Closest in time.
Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José G. C. de Souza, Shuyan Zhou, Tongshuang Wu, Graham Neubig, and André F. T. Martins · 2023
Closest in time.
Active teacher selection for reinforcement learning from human feedback, 2023
Rachel Freedman, Justin Svegliato, Kyle Wray, and Stuart Russell · 2023
Closest in time.
The Capacity for Moral Self-Correction in Large Language Models, 2023
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Mueller, Joshua Landau, Kamal Ndousse, Karina Nguyen, Liane Lovitt, Michael Sellitto, Nelson Elhage, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robert Lasenby, Robin Larson, Sam Ringer, Sandipan Kundu, Saurav Kadavath, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, Christopher Olah, Jack Clark, Samuel R. Bowman, and Jared Kaplan · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hadi Kazemi, Fariborz Taherkhani, and Nasser M. Nasrabadi · 2020
Cited alongside, same era.
Iterative Interactive Reward Learning
Pallavi Koppol, Henny Admoni, and Reid Simmons · 2020
Cited alongside, same era.
Reinforcement Learning with Augmented Data
Misha Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel, and Aravind Srinivas · 2020
Cited alongside, same era.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Cited alongside, same era.
Network Randomization: A Simple Technique for Generalization in Deep Reinforcement Learning
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 2020
Cited alongside, same era.
The Surprising Creativity of Digital Evolution: A Collection of Anecdotes from the Evolutionary Computation and Artificial Life Research Communities
Joel Lehman, Jeff Clune, Dusan Misevic, Christoph Adami, Lee Altenberg, Julie Beaulieu, Peter J. Bentley, Samuel Bernard, Guillaume Beslon, David M. Bryson, Nick Cheney, Patryk Chrabaszcz, Antoine Cully, Stephane Doncieux, Fred C. Dyer, Kai Olav Ellefsen, Robert Feldt, Stephan Fischer, Stephanie Forrest, Antoine Fŕenoy, Christian Gagńe, Leni Le Goff, Laura M. Grabowski, Babak Hodjat, Frank Hutter, Laurent Keller, Carole Knibbe, Peter Krcah, Richard E. Lenski, Hod Lipson, Robert MacCurdy, Carlos Maestre, Risto Miikkulainen, Sara Mitri, David E. Moriarty, Jean-Baptiste Mouret, Anh Nguyen, Charles Ofria, Marc Parizeau, David Parsons, Robert T. Pennock, William F. Punch, Thomas S. Ray, Marc Schoenauer, Eric Schulte, Karl Sims, Kenneth O. Stanley, François Taddei, Danesh Tarapore, Simon Thibault, Richard Watson, Westley Weimer, and Jason Yosinski · 2020
Cited alongside, same era.
A Review on Interactive Reinforcement Learning From Human Social Feedback
Jinying Lin, Zhen Ma, Randy Gomez, Keisuke Nakamura, Bo He, and Guangliang Li · 2020
Cited alongside, same era.
Closest in time.
Scaling Laws for Reward Model Overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Closest in time.
The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types
Gaurav R. Ghosal, Matthew Zurek, Daniel S. Brown, and Anca D. Dragan · 2023
Closest in time.
Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences
Lin Guan, Karthik Valmeekam, and Subbarao Kambhampati · 2023
Closest in time.
Reinforced Self-Training (ReST) for Language Modeling, 2023
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas · 2023
Closest in time.
Mastering Diverse Domains through World Models, 2023
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.
GAN-Based Interactive Reinforcement Learning from Demonstration and Human Evaluative Feedback
Jie Huang, Jiangshan Hao, Rongshun Juan, Randy Gomez, Keisuke Nakamura, and Guangliang Li · 2023
Closest in time.
Sequential Preference Ranking for Efficient Reinforcement Learning from Human Feedback
Minyoung Hwang, Gunmin Lee, Hogun Kee, Chan Woo Kim, Kyungjae Lee, and Songhwai Oh · 2023
Closest in time.
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
Beyond Reward: Offline Preference-guided Policy Optimization
Yachen Kang, Diyuan Shi, Jinxin Liu, Li He, and Donglin Wang · 2023
Closest in time.
On the Challenges and Practices of Reinforcement Learning from Real Human Feedback
Timo Kaufmann, Sarah Ball, Jacob Beck, Frauke Kreuter, and Eyke Hüllermeier · 2023
Closest in time.
Preference Transformer: Modeling Human Preferences using Transformers for RL
Changyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee · 2023
Closest in time.
OpenAssistant Conversations - Democratizing Large Language Model Alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, Shahul Es, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick · 2023
Closest in time.
Pretraining Language Models with Human Preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao, Christopher Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez · 2023
Closest in time.
Reward Design with Language Models
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh · 2023
Closest in time.
The History and Risks of Reinforcement Learning and Human Feedback, 2023
Nathan Lambert, Thomas Krendl Gilbert, and Tom Zick · 2023
Closest in time.
Aligning Text-to-Image Models using Human Feedback, 2023
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Closest in time.
Reinforcement learning with Human Feedback: Learning Dynamic Choices via Pessimism
Zihao Li, Zhuoran Yang, and Mengdi Wang · 2023
Closest in time.
Zero-shot Cross-task Preference Alignment for Offline RL via Optimal Transport
Runze Liu, Yali Du, Fengshuo Bai, Jiafei Lyu, and Xiu Li · 2023
Closest in time.
Efficient Preference-Based Reinforcement Learning Using Learned Dynamics Models
Yi Liu, Gaurav Datta, Ellen Novoseller, and Daniel S. Brown · 2023
Closest in time.
Summary of ChatGPT-Related research and perspective towards the future of large language models
Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, Zihao Wu, Lin Zhao, Dajiang Zhu, Xiang Li, Ning Qiang, Dingang Shen, Tianming Liu, and Bao Ge · 2023
Closest in time.
Aligning Human Preferences with Baseline Objectives in Reinforcement Learning
Daniel Marta, Simon Holk, Christian Pek, Jana Tumova, and Iolanda Leite · 2023
Closest in time.
Unified Learning from Demonstrations, Corrections, and Preferences during Physical Human-Robot Interaction
Shaunak A. Mehta and Dylan P. Losey · 2023
Closest in time.
Sample Efficient Reinforcement Learning from Human Feedback via Active Exploration, 2023
Viraj Mehta, Vikramjeet Das, Ojash Neopane, Yijia Dai, Ilija Bogunovic, Jeff Schneider, and Willie Neiswanger · 2023
Closest in time.
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
Katherine Metcalf, Miguel Sarabia, Natalie Mackraz, and Barry-John Theobald · 2023
Closest in time.
RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback
Yannick Metz, David Lindner, Raphaël Baur, Daniel A. Keim, and Mennatallah El-Assady · 2023
Closest in time.
Explainable Reinforcement Learning: A Survey and Comparative Review
Stephanie Milani, Nicholay Topin, Manuela Veloso, and Fei Fang · 2023
Closest in time.
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications, 2023
Abhilash Mishra · 2023
Closest in time.
On Huber’s contaminated model
Weiyan Mu and Shifeng Xiong · 2023
Closest in time.
Active Reward Learning from Online Preferences
Vivek Myers, Erdem Bıyık, and Dorsa Sadigh · 2023
Closest in time.
DIP-RL: Demonstration-Inferred Preference Learning in Minecraft
Ellen Novoseller, Vinicius G. Goecks, David Watkins, Josh Miller, and Nicholas R. Waytowich · 2023
Closest in time.
GPT-4 Technical Report
OpenAI · 2023
Closest in time.
Epistemic Neural Networks
Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, and Benjamin Van Roy · 2023
Closest in time.
Tuning Computer Vision Models With Task Rewards
André Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer, and Xiaohua Zhai · 2023
Closest in time.
Learning Rewards to Optimize Global Performance Metrics in Deep Reinforcement Learning
Junqi Qian, Paul Weng, and Chenmien Tan · 2023
Closest in time.
A Survey on Explainable Reinforcement Learning: Concepts, Algorithms, Challenges, 2023
Yunpeng Qing, Shunyu Liu, Jie Song, Huiqiong Wang, and Mingli Song · 2023
Closest in time.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn · 2023
Closest in time.
Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2023
Closest in time.
Dueling RL: Reinforcement Learning with Trajectory Preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee · 2023
Closest in time.
Reinforcement Learning from Human Feedback: Progress and Challenges, 2023
John Schulman · 2023
Closest in time.
Contextual Bandits and Imitation Learning with Preference-Based Active Queries
Ayush Sekhari, Karthik Sridharan, Wen Sun, and Runzhe Wu · 2023
Closest in time.
Benchmarks and Algorithms for Offline Preference-Based Reward Learning
Daniel Shin, Anca Dragan, and Daniel S. Brown · 2023
Closest in time.
Fairness in Preference-based Reinforcement Learning
Umer Siddique, Abhinav Sinha, and Yongcan Cao · 2023
Closest in time.
Misspecification in Inverse Reinforcement Learning
Joar Max Viktor Skalse and Alessandro Abate · 2023
Closest in time.
Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2023
Closest in time.
Reward Collapse in Aligning Large Language Models: A Prompt-Aware Approach to Preference Rankings
Ziang Song, Tianle Cai, Jason D. Lee, and Weijie J. Su · 2023
Closest in time.
Causal Confusion and Reward Misidentification in Preference-Based Reward Learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca Dragan, and Daniel S. Brown · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Closest in time.
Data Driven Reward Initialization for Preference based Reinforcement Learning
Mudit Verma and Subbarao Kambhampati · 2023
Closest in time.
A State Augmentation based approach to Reinforcement Learning from Human Preferences
Mudit Verma and Subbarao Kambhampati · 2023
Closest in time.
Exploiting Unlabeled Data for Feedback Efficient Human Preference based Reinforcement Learning
Mudit Verma, Siddhant Bhambri, and Subbarao Kambhampati · 2023
Closest in time.
Self-Instruct: Aligning Language Models with Self-Generated Instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Closest in time.
Is RLHF More Difficult than Standard RL? A Theoretical Perspective
Yuanhao Wang, Qinghua Liu, and Chi Jin · 2023
Closest in time.
Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Tianhao Wu, Banghua Zhu, Ruoyu Zhang, Zhaojin Wen, Kannan Ramchandran, and Jiantao Jiao · 2023
Closest in time.
ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong · 2023
Closest in time.
PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User Engagement
Wanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun, Shuchang Liu, Dong Zheng, Peng Jiang, Kun Gai, and Bo An · 2023
Closest in time.
Scaling Robot Learning with Semantically Imagined Experience
Tianhe Yu, Ted Xiao, Jonathan Tompson, Austin Stone, Su Wang, Anthony Brohan, Jaspiar Singh, Clayton Tan, Dee M, Jodilyn Peralta, Karol Hausman, Brian Ichter, and Fei Xia · 2023
Closest in time.
RRHF: Rank Responses to Align Language Models with Human Feedback
Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Closest in time.
Learning state importance for preference-based reinforcement learning
Guoxi Zhang and Hisashi Kashima · 2023
Closest in time.
SLiC-HF: Sequence Likelihood Calibration with Human Feedback, 2023
Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, and Peter J. Liu · 2023
Closest in time.
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Closest in time.
A General Theoretical Paradigm to Understand Learning from Human Preferences
Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello · 2024
Closest in time.
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
Hritik Bansal, John Dang, and Aditya Grover · 2024
Closest in time.
Learning Interpretable Models of Aircraft Handling Behaviour by Reinforcement Learning from Human Feedback
Tom Bewley, Jonathan Lawry, and Arthur Richards · 2024
Closest in time.
Batch Active Learning of Reward Functions from Human Preferences
Erdem Bıyık, Nima Anari, and Dorsa Sadigh · 2024
Closest in time.
Suppressing Pink Elephants with Direct Principle Feedback, 2024
Louis Castricato, Nathan Lile, Suraj Anand, Hailey Schoelkopf, Siddharth Verma, and Stella Biderman · 2024
Closest in time.
MaxMin-RLHF: Towards Equitable Alignment of Large Language Models with Diverse Human Preferences
Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Furong Huang, Dinesh Manocha, Amrit Bedi, and Mengdi Wang · 2024
Closest in time.
Dense Reward for Free in Reinforcement Learning from Human Feedback
Alex James Chan, Hao Sun, Samuel Holt, and Mihaela van der Schaar · 2024
Closest in time.
Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback
Yu Chen, Yihan Du, Pihe Hu, Siwei Wang, Desheng Wu, and Longbo Huang · 2024
Closest in time.
Crowd-PrefRL: Preference-Based Reward Learning from Crowds, 2024
David Chhan, Ellen Novoseller, and Vernon J. Lawhern · 2024
Closest in time.
MusicRL: Aligning Music Generation to Human Preferences
Geoffrey Cideron, Sertan Girgin, Mauro Verzetti, Damien Vincent, Matej Kastelic, Zalán Borsos, Brian Mcwilliams, Victor Ungureanu, Olivier Bachem, Olivier Pietquin, Matthieu Geist, Leonard Hussenot, Neil Zeghidour, and Andrea Agostinelli · 2024
Closest in time.
Position: Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mosse, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, and William S. Zwicker · 2024
Closest in time.
Reward Model Ensembles Help Mitigate Overoptimization
Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger · 2024
Closest in time.
Mapping Social Choice Theory to RLHF
Jessica Dai and Eve Fleisig · 2024
Closest in time.
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
A density estimation perspective on learning from pairwise human preferences
Vincent Dumoulin, Daniel D. Johnson, Pablo Samuel Castro, Hugo Larochelle, and Yann Dauphin · 2024
Closest in time.
Efficient Exploration for LLMs
Vikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao, and Benjamin Van Roy · 2024
Closest in time.
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Jacob Eisenstein, Chirag Nagpal, Alekh Agarwal, Ahmad Beirami, Alexander Nicholas D’Amour, Krishnamurthy Dj Dvijotham, Adam Fisch, Katherine A. Heller, Stephen Robert Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant · 2024
Closest in time.
A Generalized Acquisition Function for Preference-based Reward Learning
Evan Ellis, Gaurav R. Ghosal, Stuart J. Russell, Anca Dragan, and Erdem Bıyık · 2024
Closest in time.
A survey on interpretable reinforcement learning
Claire Glanois, Paul Weng, Matthieu Zimmer, Dong Li, Tianpei Yang, Jianye Hao, and Wulong Liu · 2024
Closest in time.
The False Promise of Imitating Proprietary Language Models
Arnav Gudibande, Eric Wallace, Charlie Victor Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2024
Closest in time.
Contrastive Preference Learning: Learning from Human Feedback without Reinforcement Learning
Joey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn, Scott Niekum, W. Bradley Knox, and Dorsa Sadigh · 2024
Closest in time.
Human Feedback is not Gold Standard
Tom Hosking, Phil Blunsom, and Max Bartolo · 2024
Closest in time.
The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
Shengyi Huang, Michael Noukhovitch, Arian Hosseini, Kashif Rasul, Weixun Wang, and Lewis Tunstall · 2024
Closest in time.
Can Differentiable Decision Trees Enable Interpretable Reward Learning from Human Feedback?
Akansha Kalra and Daniel S Brown · 2024
Closest in time.
Motif: Intrinsic Motivation from Artificial Intelligence Feedback
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff · 2024
Closest in time.
Models of human preference for learning reward functions
W. Bradley Knox, Stephane Hatgis-Kessell, Serena Booth, Scott Niekum, Peter Stone, and Alessandro G. Allievi · 2024
Closest in time.
RewardBench: Evaluating Reward Models for Language Modeling, 2024
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, L. J. Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Closest in time.
Generative Judge for Evaluating Alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu · 2024
Closest in time.
RLIF: Interactive Imitation Learning as Reinforcement Learning
Jianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma, and Sergey Levine · 2024
Closest in time.
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He · 2024
Closest in time.
Confronting Reward Model Overoptimization with Constrained RLHF
Ted Moskovitz, Aaditya K. Singh, D. J. Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca Dragan, and Stephen Marcus McAleer · 2024
Closest in time.
Nash Learning from Human Feedback
Remi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Zhaohan Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Côme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, and Bilal Piot · 2024
Closest in time.
The Alignment Problem from a Deep Learning Perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2024
Closest in time.
Introducing ChatGPT, 2022
OpenAI · 2024
Closest in time.
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
Chanwoo Park, Mingyang Liu, Dingwen Kong, Kaiqing Zhang, and Asuman E. Ozdaglar · 2024
Closest in time.
Towards interactive reinforcement learning with intrinsic feedback
Benjamin Poole and Minwoo Lee · 2024
Closest in time.
WARM: On the Benefits of Weight Averaged Reward Models
Alexandre Rame, Nino Vieillard, Leonard Hussenot, Robert Dadashi, Geoffrey Cideron, Olivier Bachem, and Johan Ferret · 2024
Closest in time.
Human-in-the-Loop Reinforcement Learning: A Survey and Position on Requirements, Challenges, and Opportunities
Carl Orge Retzlaff, Srijita Das, Christabel Wayllace, Payam Mousavi, Mohammad Afshari, Tianpei Yang, Anna Saranti, Alessa Angerschmid, Matthew E. Taylor, and Andreas Holzinger · 2024
Closest in time.
The Trickle-down Impact of Reward Inconsistency on RLHF
Lingfeng Shen, Sihao Chen, Linfeng Song, Lifeng Jin, Baolin Peng, Haitao Mi, Daniel Khashabi, and Dong Yu · 2024
Closest in time.
Understanding Hidden Context in Preference Learning: Consequences for RLHF
Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell · 2024
Closest in time.
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
Joar Max Viktor Skalse and Alessandro Abate · 2024
Closest in time.
STARC: A General Framework For Quantifying Differences Between Reward Functions
Joar Max Viktor Skalse, Lucy Farnik, Sumeet Ramesh Motwani, Erik Jenner, Adam Gleave, and Alessandro Abate · 2024
Closest in time.
Preference Ranking Optimization for Human Alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang · 2024
Closest in time.
Position: A Roadmap to Pluralistic Alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L. Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi · 2024
Closest in time.
SALMON: Self-Alignment with Instructable Reward Models
Zhiqing Sun, Yikang Shen, Hongxin Zhang, Qinhong Zhou, Zhenfang Chen, David Daniel Cox, Yiming Yang, and Chuang Gan · 2024
Closest in time.
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu, and Alekh Agarwal · 2024
Closest in time.
Zeroth-Order Optimization Meets Human Feedback: Provable Learning via Ranking Oracles
Zhiwei Tang, Dmitry Rybin, and Tsung-Hui Chang · 2024
Closest in time.
Hindsight PRIORs for Reward Learning from Human Preferences
Mudit Verma and Katherine Metcalf · 2024
Closest in time.
A comprehensive survey on deep active learning in medical image analysis
Haoran Wang, Qiuye Jin, Shiman Li, Siyu Liu, Manning Wang, and Zhijian Song · 2024
Closest in time.
HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Scowcroft, Neel Kant, Aidan Swope, and Oleksii Kuchaiev · 2024
Closest in time.
Privately Aligning Language Models with Reinforcement Learning
Fan Wu, Huseyin A. Inan, Arturs Backurs, Varun Chandrasekaran, Janardhan Kulkarni, and Robert Sim · 2024
Closest in time.
Making RL with Preference-based Feedback Efficient via Randomization
Runzhe Wu and Wen Sun · 2024
Closest in time.
Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Tianbao Xie, Siheng Zhao, Chen Henry Wu, Yitao Liu, Qian Luo, Victor Zhong, Yanchao Yang, and Tao Yu · 2024
Closest in time.
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang · 2024
Closest in time.
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Shusheng Xu, Wei Fu, Jiaxuan Gao, Wenjie Ye, Weilin Liu, Zhiyu Mei, Guangju Wang, Chao Yu, and Yi Wu · 2024
Closest in time.
Reinforcement Learning from Diverse Human Preferences
Wanqi Xue, Bo An, Shuicheng Yan, and Zhongwen Xu · 2024
Closest in time.
Capability or Alignment? Respect the LLM Base Model’s Capability During Alignment, 2024
Jingfeng Yang · 2024
Closest in time.
Chenlu Ye, Wei Xiong, Yuheng Zhang, Nan Jiang, and Tong Zhang · 2024
Closest in time.
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
Yifu Yuan, Jianye Hao, Yi Ma, Zibin Dong, Hebin Liang, Jinyi Liu, Zhixin Feng, Kai Zhao, and Yan Zheng · 2024
Closest in time.
CPPO: Continual Learning for Reinforcement Learning with Human Feedback
Han Zhang, Yu Lei, Lin Gui, Min Yang, Yulan He, Hui Wang, and Ruifeng Xu · 2024
Closest in time.
Improving Generalization of Alignment with Human Preferences through Group Invariant Learning
Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.
Panacea: Pareto Alignment via Preference Adaptation for LLMs
Yifan Zhong, Chengdong Ma, Xiaoyuan Zhang, Ziran Yang, Haojun Chen, Qingfu Zhang, Siyuan Qi, and Yaodong Yang · 2024
Closest in time.
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Yifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine, and Aviral Kumar · 2024
Closest in time.
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2024
Closest in time.
Self-Improving Robust Preference Optimization
Eugene Choi, Arash Ahmadian, Matthieu Geist, Olivier Pietquin, and Mohammad Gheshlaghi Azar · 2025
Closest in time.