Fetching the paper…
Reading the bibliography…
Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs).
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019 · 1907
Earlier work this paper cites.
Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019 · 1908
Earlier work this paper cites.
Fine-Tuning Language Models from Human Preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
BERT has a Moral Compass: Improvements of ethical and moral values of machines
Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin Rothkopf, and Kristian Kersting. 2019 · 1912
Earlier work this paper cites.
Make The Most of Prior Data: A Solution for Interactive Text Summarization with Preference Feedback
Duy-Hung Nguyen, Nguyen Viet Dung Nghiem, Bao-Sinh Nguyen, Dung Tien Tien Le, Shahab Sabahi, Minh-Tien Nguyen, and Hung Le. 2022 · 1930
Earlier work this paper cites.
Intransitivity of preferences
Amos Tversky. 1969 · 1969
Earlier work this paper cites.
Introduction to the work of Marcel Mauss
Claude Lévi-Strauss. 1987 · 1987
Earlier work this paper cites.
The biasing effects of scale-checking styles on response to a likert scale
Hershey H Friedman, Paul J Herskovitz, and Simcha Pollack. 1994 · 1994
Earlier work this paper cites.
Long Short-term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A theory of nonseparable preferences in survey responses
Dean Lacy. 2001 · 2001
Earlier work this paper cites.
On populist reason
Ernesto Laclau. 2005 · 2005
Earlier work this paper cites.
Donald Martin Jr., Vinodkumar Prabhakaran, Jill Kuhlberg, Andrew Smart, and William S. Isaac. 2020 · 2005
Earlier work this paper cites.
Making sense of social research: How useful is the hawthorne effect?
Mecca Chiesa and Sandy Hobbs. 2008 · 2008
Earlier work this paper cites.
Aligning AI With Shared Human Values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020 · 2008
Earlier work this paper cites.
In search of homo economicus: Cognitive noise and the role of emotion in preference consistency
Leonard Lee, On Amir, and Dan Ariely. 2009 · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul Christiano. 2020 · 2009
Earlier work this paper cites.
Affect in web interfaces: A study of the impacts of web page visual complexity and order
Liqiong Deng and Marshall Scott Poole. 2010 · 2010
Earlier work this paper cites.
Recipes for Safety in Open-domain Chatbots
Jing Xu, Da Ju, Margaret Li, Y.-Lan Boureau, Jason Weston, and Emily Dinan. 2021b · 2010
Earlier work this paper cites.
" yours is better!" participant response bias in hci
Nicola Dell, Vidya Vaidyanathan, Indrani Medhi, Edward Cutrell, and William Thies. 2012 · 2012
Earlier work this paper cites.
Comparative analysis of verbal alignment in human-human and human-agent interactions
Sabrina Campano, Jessica Durand, and Chloé Clavel. 2014 · 2014
Earlier work this paper cites.
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. 2015 · 2015
Earlier work this paper cites.
Stability of experimental and survey measures of risk, time, and social preferences: A review and some new results
Yating Chuang and Laura Schechter. 2015 · 2015
Earlier work this paper cites.
Response option configuration of online administered likert scales
Hotaka Maeda. 2015 · 2015
Earlier work this paper cites.
Motivating Personality-aware Machine Translation
Shachar Mirkin, Scott Nowson, Caroline Brun, and Julien Perez. 2015 · 2015
Earlier work this paper cites.
Predicting human similarity judgments with distributional models: The value of word associations
Simon De Deyne, Amy Perfors, and Daniel J Navarro. 2016 · 2016
Earlier work this paper cites.
You get who you pay for: The impact of incentives on participation bias
Gary Hsieh and Rafał Kocielnik. 2016 · 2016
Earlier work this paper cites.
Deep Reinforcement Learning for Dialogue Generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016 · 2016
Earlier work this paper cites.
Personalizing a dialogue system with transfer reinforcement learning
Kaixiang Mo, Shuangyin Li, Yu Zhang, Jiajun Li, and Qiang Yang. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access
Bhuwan Dhingra, Lihong Li, Xiujun Li, Jianfeng Gao, Yun-Nung Chen, Faisal Ahmed, and Li Deng. 2017 · 2017
Earlier work this paper cites.
Automatic Measures to Characterise Verbal Alignment in Human-Agent Interaction
Guillaume Dubuisson Duplessis, Chloé Clavel, and Frédéric Landragin. 2017 · 2017
Earlier work this paper cites.
Controlling Linguistic Style Aspects in Neural Language Generation
Jessica Ficler and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
Personality, Values, Culture
Ronald Fischer. 2017 · 2017
Earlier work this paper cites.
Bandit Structured Prediction for Neural Sequence-to-Sequence Learning
Julia Kreutzer, Artem Sokolov, and Stefan Riezler. 2017 · 2017
Earlier work this paper cites.
Counterfactual Learning from Bandit Feedback under Deterministic Logging : A Case Study in Statistical Machine Translation
Carolin Lawrence, Artem Sokolov, and Stefan Riezler. 2017 · 2017
Earlier work this paper cites.
Acquiring Background Knowledge to Improve Moral Value Prediction
Ying Lin, Joe Hoover, Morteza Dehghani, Marlon Mooijman, and Heng Ji. 2017 · 2017
Earlier work this paper cites.
A Societal Sentiment Analysis: Predicting the Values and Ethics of Individuals by Analysing Social Media Content
Tushar Maheshwari, Aishwarya N. Reganti, Samiksha Gupta, Anupam Jamatia, Upendra Kumar, Björn Gambäck, and Amitava Das. 2017 · 2017
Earlier work this paper cites.
Reinforcement Learning for Bandit Neural Machine Translation with Simulated Human Feedback
Khanh Nguyen, Hal Daumé III, and Jordan Boyd-Graber. 2017 · 2017
Earlier work this paper cites.
Personalized Machine Translation: Preserving Original Author Traits
Ella Rabinovich, Raj Nath Patel, Shachar Mirkin, Lucia Specia, and Shuly Wintner. 2017 · 2017
Earlier work this paper cites.
A Computational Model of Human Preferences for Pronoun Resolution
Olga Seminck and Pascal Amsili. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Predicting Users’ Negative Feedbacks in Multi-Turn Human-Computer Dialogues
Xin Wang, Jianan Wang, Yuanchao Liu, Xiaolong Wang, Zhuoran Wang, and Baoxun Wang. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
APRIL: Interactively learning to summarise by combining active preference learning and reinforcement learning
Yang Gao, Christian M. Meyer, and Iryna Gurevych. 2018 · 2018
Earlier work this paper cites.
Reliability and Learnability of Human Bandit Feedback for Sequence-to-Sequence Reinforcement Learning
Julia Kreutzer, Joshua Uyheng, and Stefan Riezler. 2018 · 2018
Earlier work this paper cites.
Improving a Neural Semantic Parser by Counterfactual Learning from Human Bandit Feedback
Carolin Lawrence and Stefan Riezler. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations. arxiv 2018
ME Peters, M Neumann, M Iyyer, M Gardner, C Clark, K Lee, and L Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Personalizing Dialogue Agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Cited alongside, same era.
Better Rewards Yield Better Summaries: Learning to Summarise Without References
Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Bias and fairness in natural language processing
Kai-Wei Chang, Vinodkumar Prabhakaran, and Vicente Ordonez. 2019 · 2019
Cited alongside, same era.
Reinforcement learning for personalized dialogue management
Floris Den Hengst, Mark Hoogendoorn, Frank Van Harmelen, and Joost Bosman. 2019 · 2019
Cited alongside, same era.
Enabling Classifiers to Make Judgements Explicitly Aligned with Human Values
Yejin Bang, Tiezheng Yu, Andrea Madotto, Zhaojiang Lin, Mona Diab, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Power to the People? Opportunities and Challenges for Participatory AI
Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Diaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022 · 2022
Later among the works it cites.
Measuring Progress on Scalable Oversight for Large Language Models
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
Robust preference learning for storytelling via contrastive reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Do RNNs learn human-like abstract word order preferences?
Richard Futrell and Roger P. Levy. 2019 · 2019
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
Learning from Dialogue after Deployment: Feed Yourself, Chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Semantics Derived Automatically from Language Corpora Contain Human-like Moral Choices
Sophie Jentzsch, Patrick Schramowski, Constantin Rothkopf, and Kristian Kersting. 2019 · 2019
Cited alongside, same era.
The global landscape of ai ethics guidelines
Anna Jobin, Marcello Ienca, and Effy Vayena. 2019 · 2019
Cited alongside, same era.
Multilingual Entity, Relation, Event and Human Value Extraction
Manling Li, Ying Lin, Joseph Hoover, Spencer Whitehead, Clare Voss, Morteza Dehghani, and Heng Ji. 2019 · 2019
Cited alongside, same era.
Louis Castricato, Alexander Havrilla, Shahbuland Matiana, Michael Pieler, Anbang Ye, Ian Yang, Spencer Frazier, and Mark Riedl. 2022 · 2022
Later among the works it cites.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022 · 2022
Later among the works it cites.
Toward personalized answer generation in e-commerce via multi-perspective preference modeling
Yang Deng, Yaliang Li, Wenxuan Zhang, Bolin Ding, and Wai Lam. 2022 · 2022
Later among the works it cites.
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022 · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Jury learning: Integrating dissenting voices into machine learning models
Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022 · 2022
Later among the works it cites.
Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022 · 2022
Later among the works it cites.
Learning to Adapt Domain Shifts of Moral Values via Instance Weighting
Xiaolei Huang, Alexandra Wormley, and Adam Cohen. 2022 · 2022
Later among the works it cites.
Survey of Hallucination in Natural Language Generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022 · 2022
Later among the works it cites.
Can Machines Learn Morality? The Delphi Experiment
Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi. 2022 · 2022
Later among the works it cites.
When to Make Exceptions: Exploring Language Models as Accounts of Human Moral Judgment
Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf. 2022 · 2022
Later among the works it cites.
Da Ju, Jing Xu, Y.-Lan Boureau, and Jason Weston. 2022 · 2022
Later among the works it cites.
Identifying the Human Values behind Arguments
Johannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, and Benno Stein. 2022 · 2022
Later among the works it cites.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022 · 2022
Later among the works it cites.
Aligning Generative Language Models with Human Values
Ruibo Liu, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. 2022 · 2022
Later among the works it cites.
Towards Boosting the Open-Domain Chatbot with Human Feedback
Hua Lu, Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2022 · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell-Gillingham, Geoffrey Irving, and Nat McAleese. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
"More Than Words": Linking Music Preferences and Moral Values Through Lyrics
Vjosa Preniqi, Kyriaki Kalimeri, and Charalampos Saitis. 2022 · 2022
Later among the works it cites.
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, and Chandra Bhagavatula. 2022 · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective nlp tasks
Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
Training Language Models with Language Feedback
Jérémy Scheurer, Jon Ander Campos, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2022 · 2022
Later among the works it cites.
"I’m sorry to hear that": Finding bias in language models with a holistic descriptor dataset
Eric Michael Smith, Melissa Hall Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022 · 2022
Later among the works it cites.
LaMDA: Language Models for Dialog Applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022 · 2022
Later among the works it cites.
Self-Instruct: Aligning Language Model with Self Generated Instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022 · 2022
Later among the works it cites.
Information loss and bias in likert survey responses
J Christopher Westland. 2022 · 2022
Later among the works it cites.
Jing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora, Y.-Lan Boureau, and Jason Weston. 2022 · 2022
Later among the works it cites.
The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
Caleb Ziems, Jane Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022 · 2022
Later among the works it cites.
Participatory Design of AI Systems: Opportunities and Challenges Across Diverse Users, Relationships, and Application Domains
Douglas Zytko, Pamela J. Wisniewski, Shion Guha, Eric P. S. Baumer, and Min Kyung Lee. 2022 · 2022
Later among the works it cites.
Assessing Language Model Deployment with Risk Cards
Leon Derczynski, Hannah Rose Kirk, Vidhisha Balachandran, Sachin Kumar, Yulia Tsvetkov, M. R. Leiser, and Saif Mohammad. 2023 · 2023
Closest in time.
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. 2023 · 2023
Closest in time.
Human Feedback is not Gold Standard
Tom Hosking, Phil Blunsom, and Max Bartolo. 2023 · 2023
Closest in time.
Pretraining Language Models with Human Preferences
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez. 2023 · 2023
Closest in time.
The touché23-ValueEval dataset for identifying human values behind arguments
Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Barriere Valentin, Doratossadat Dastgheib, Omid Ghahroodi, Mohammad Ali Sadraei, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, and Benno Stein. 2023 · 2023
Closest in time.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023 · 2023
Closest in time.
Preference Ranking Optimization for Human Alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang. 2023 · 2023
Closest in time.
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. 2023 · 2023
Closest in time.
LIMA: Less Is More for Alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023 · 2023
Closest in time.
Personalized machine translation: Predicting translational preferences
Shachar Mirkin and Jean-Luc Meunier. 2015 · 2025
Closest in time.
Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems
Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018 · 2069
Closest in time.