Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) have become central tools in various fields, they often provide inaccurate or false information.
Fine-Tuning Language Models from Human Preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Intention—behavior relations: a conceptual and empirical review
Paschal Sheeran. 2002 · 2002
Earlier work this paper cites.
Influence: The psychology of persuasion , volume 55
Robert B Cialdini. 2007 · 2007
Earlier work this paper cites.
Running experiments on amazon mechanical turk
Gabriele Paolacci, Jesse Chandler, and Panagiotis G Ipeirotis. 2010 · 2010
Earlier work this paper cites.
Why we don’t “just do it” understanding the intention-behavior gap in lifestyle medicine
Mark D Faries. 2016 · 2016
Earlier work this paper cites.
Can personality close the intention-behavior gap for healthy eating? An examination with the HEXACO personality traits
Lauren A Monds, Carolyn MacCann, Barbara A Mullan, Cara Wong, Jemma Todd, and Richard D Roberts. 2016 · 2016
Earlier work this paper cites.
The spread of true and false news online
Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018 · 2018
Earlier work this paper cites.
Fighting misinformation on social media using crowdsourced judgments of news source quality
Gordon Pennycook and David G Rand. 2019 · 2019
Earlier work this paper cites.
Why do people spread false information online? The effects of message and viewer characteristics on self-reported likelihood of sharing social media disinformation
Tom Buchanan. 2020 · 2020
Cited alongside, same era.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al. 2022 · 2022
Cited alongside, same era.
The Internal State of an LLM Knows When It’s Lying
Amos Azaria and Tom M. Mitchell. 2023 · 2023
Cited alongside, same era.
Can AI language models replace human participants?
Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. 2023 · 2023
Cited alongside, same era.
Large language models as simulated economic agents: What can we learn from homo silicus?
Horton, John J. 2023 · 2023
Cited alongside, same era.
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
Hannah Kirk, Andrew M. Bean, Bertie Vidgen, Paul Röttger, and Scott Hale. 2023 · 2023
Later among the works it cites.
Self-Refine: Iterative Refinement with Self-Feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Later among the works it cites.
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023 · 2023
Later among the works it cites.
A systematic literature review of user trust in AI-enabled systems: An HCI perspective
Tita Alissa Bach, Amna Khan, Harry Hallock, Gabriela Beltrão, and Sonia Sousa. 2024 · 2024
Closest in time.
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2023 · 2023
Cited alongside, same era.
A survey of reinforcement learning from human feedback
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2023 · 2023
Cited alongside, same era.
ChatGPT 3.5 (June 16 version) [Large language model]. https://openai.com
OpenAI. 2024a
Cited in the paper.
ChatGPT 4 (November 28 version) [Large language model]. https://openai.com
OpenAI. 2024b
Cited in the paper.
Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, and Bruno Castro da Silva. 2024 · 2024
Closest in time.
Understanding the Effects of RLHF on LLM Generalisation and Diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. 2024 · 2024
Closest in time.
BOND: Aligning LLMs with Best-of-N Distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ramé, Bobak Shahriari, Sarah Perrin, Abe Friesen, Geoffrey Cideron, Sertan Girgin, Piotr Stanczyk, Andrea Michi, Danila Sinopalnikov, Sabela Ramos, Amélie Héliou, Aliaksei Severyn, Matt Hoffman, Nikola Momchev, and Olivier Bachem. 2024 · 2024
Closest in time.