Fetching the paper…
Reading the bibliography…
LLMs are aligned to follow input instructions by learning which of two responses users prefer for a prompt.
The process of asking questions
Robert S Taylor. 1962 · 1962
Earlier work this paper cites.
Collected papers of charles sanders peirce , volume 5
Charles Sanders Peirce. 1974 · 1974
Earlier work this paper cites.
The measurement of interrater agreement
Joseph L Fleiss, Bruce Levin, Myunghee Cho Paik, et al. 1981 · 1981
Earlier work this paper cites.
Varieties of confirmation bias
Joshua Klayman. 1995 · 1995
Earlier work this paper cites.
A structural approach to selection bias
Miguel A Hernán, Sonia Hernández-Díaz, and James M Robins. 2004 · 2004
Earlier work this paper cites.
Making it personal: How personalization affects trust over time
Catharina M Serino, Christopher P Furner, and Cindi Smatt. 2005 · 2005
Earlier work this paper cites.
Challenging the long tail recommendation
Hongzhi Yin, Bin Cui, Jing Li, Junjie Yao, and Chen Chen. 2012 · 2012
Earlier work this paper cites.
TL;DR: Mining Reddit to learn automatic summarization
Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein. 2017 · 2017
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
A literature review of quantitative persona creation
Joni Salminen, Kathleen Guan, Soon-Gyo Jung, Shammur A. Chowdhury, and Bernard J. Jansen. 2020 · 2020
Earlier work this paper cites.
A systematic review of research on personalized learning: Personalized by whom, to what, how, and for what purpose (s)?
Matthew L Bernacki, Meghan J Greene, and Nikki G Lobczowski. 2021 · 2021
Earlier work this paper cites.
Learning event graph knowledge for abductive reasoning
Li Du, Xiao Ding, Ting Liu, and Bing Qin. 2021 · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. 2021 · 2021
Earlier work this paper cites.
Causal direction of data collection matters: Implications of causal and anticausal learning for NLP
Zhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, and Bernhard Schoelkopf. 2021 · 2021
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al. 2022 · 2022
Earlier work this paper cites.
Boosting natural language generation from instructions with meta-learning
Budhaditya Deb, Ahmed Hassan Awadallah, and Guoqing Zheng. 2022 · 2022
Earlier work this paper cites.
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
An empirical survey of the effectiveness of debiasing techniques for pre-trained language models
Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
The role of personalization in the user experience, preferences and engagement with virtual reality environments for relaxation
Susanna Pardini, Silvia Gabrielli, Marco Dianti, Caterina Novara, Gesualdo M Zucco, Ornella Mich, and Stefano Forti. 2022 · 2022
Earlier work this paper cites.
ColBERTv2: Effective and efficient retrieval via lightweight late interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022 · 2022
Earlier work this paper cites.
Training data is more valuable than you think: A simple and effective method by retrieving from training data
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, and Michael Zeng. 2022 · 2022
Earlier work this paper cites.
Confirmation bias and the persistence of misinformation on climate change
Yanmengqian Zhou and Lijiang Shen. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Toxicity in chatgpt: Analyzing persona-assigned language models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Natural language decompositions of implicit content enable better text representations
Alexander Hoyle, Rupak Sarkar, Pranav Goel, and Philip Resnik. 2023 · 2023
Cited alongside, same era.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023 · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023 · 2023
Cited alongside, same era.
Ai alignment: A comprehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. 2023 · 2023
Cited alongside, same era.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Evaluating large language model biases in persona-steered generation
Andy Liu, Mona Diab, and Daniel Fried. 2024a · 2024
Later among the works it cites.
Contextualized evaluations: Taking the guesswork out of language model evaluations
Chaitanya Malaviya, Joseph Chee Chang, Dan Roth, Mohit Iyyer, Mark Yatskar, and Kyle Lo. 2024 · 2024
Later among the works it cites.
Beyond the binary: Capturing diverse preferences with reward regularization
Vishakh Padmakumar, Chuanyang Jin, Hannah Rose Kirk, and He He. 2024 · 2024
Later among the works it cites.
PocketLLM: Enabling on-device fine-tuning for personalized LLMs
Dan Peng, Zhihui Fu, and Jun Wang. 2024 · 2024
Later among the works it cites.
Improving context-aware preference modeling for language models
Silviu Pitis, Ziang Xiao, Nicolas Le Roux, and Alessandro Sordoni. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What makes it ok to set a fire? iterative self-distillation of contexts and rationales for disambiguating defeasible social and moral situations
Kavel Rao, Liwei Jiang, Valentina Pyatkin, Yuling Gu, Niket Tandon, Nouha Dziri, Faeze Brahman, and Yejin Choi. 2023 · 2023
Cited alongside, same era.
Can authorship representation learning capture stylistic features?
Andrew Wang, Cristina Aggazzotti, Rebecca Kotula, Rafael Rivera Soto, Marcus Bishop, and Nicholas Andrews. 2023 · 2023
Cited alongside, same era.
Abductive commonsense reasoning exploiting mutually exclusive explanations
Wenting Zhao, Justin Chiu, Claire Cardie, and Alexander Rush. 2023 · 2023
Cited alongside, same era.
Goal driven discovery of distributional differences via language descriptions
Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Dan Klein, and Jacob Steinhardt. 2023 · 2023
Cited alongside, same era.
The illusion of artificial inclusion
William Agnew, A. Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R. McKee. 2024 · 2024
Cited alongside, same era.
Meet claude
Anthropic. 2023 · 2024
Cited alongside, same era.
It’s not easy being wrong: Large language models struggle with process of elimination reasoning
Nishant Balepur, Shramay Palta, and Rachel Rudinger. 2024a · 2024
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Later among the works it cites.
LaMP: When large language models meet personalization
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024 · 2024
Later among the works it cites.
The prompt report: A systematic survey of prompting techniques
Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al. 2024 · 2024
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2024 · 2024
Later among the works it cites.
KARL: Knowledge-aware retrieval and representations aid retention and learning in students
Matthew Shu, Nishant Balepur, Shi Feng, and Jordan Lee Boyd-Graber. 2024 · 2024
Later among the works it cites.
Position: A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L. Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024 · 2024
Later among the works it cites.
Rlvf: learning from verbal feedback without overgeneralization
Moritz Stephan, Alexander Khazatsky, Eric Mitchell, Annie S Chen, Sheryl Hsu, Archit Sharma, and Chelsea Finn. 2024 · 2024
Later among the works it cites.
Two tales of persona in LLMs: A survey of role-playing and personalization
Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Wei-Lin Chen, Chao-Wei Huang, Yu Meng, and Yun-Nung Chen. 2024 · 2024
Later among the works it cites.
Do-not-answer: Evaluating safeguards in llms
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2024c · 2024
Later among the works it cites.
Re-reading improves reasoning in large language models
Xiaohan Xu, Chongyang Tao, Tao Shen, Can Xu, Hongbo Xu, Guodong Long, Jian-Guang Lou, and Shuai Ma. 2024 · 2024
Later among the works it cites.
Bayesian reward models for LLM alignment
Adam X. Yang, Maxime Robeyns, Thomas Coste, Jun Wang, Haitham Bou Ammar, and Laurence Aitchison. 2024 · 2024
Later among the works it cites.
On diversified preferences of large language model alignment
Dun Zeng, Yong Dai, Pengyu Cheng, Longyue Wang, Tianhao Hu, Wanshun Chen, Nan Du, and Zenglin Xu. 2024 · 2024
Later among the works it cites.
Fair abstractive summarization of diverse perspectives
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, and Rui Zhang. 2024b · 2024
Later among the works it cites.
UNcommonsense reasoning: Abductive reasoning about uncommon situations
Wenting Zhao, Justin Chiu, Jena Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Li, and Alane Suhr. 2024 · 2024
Later among the works it cites.
When ”a helpful assistant” is not really helpful: Personas in system prompts do not improve performances of large language models
Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens. 2024b · 2024
Later among the works it cites.
Reverse question answering: Can an LLM write a question so hard (or bad) that it can‘t answer?
Nishant Balepur, Feng Gu, Abhilasha Ravichander, Shi Feng, Jordan Lee Boyd-Graber, and Rachel Rudinger. 2025a · 2025
Closest in time.
MoDS: Moderating a mixture of document speakers to summarize debatable queries in document collections
Nishant Balepur, Alexa Siu, Nedim Lipka, Franck Dernoncourt, Tong Sun, Jordan Lee Boyd-Graber, and Puneet Mathur. 2025b · 2025
Closest in time.
Language models predict empathy gaps between social in-groups and out-groups
Yu Hou, Hal Daumé Iii, and Rachel Rudinger. 2025 · 2025
Closest in time.
Improving llm personas via rationalization with psychological scaffolds
Brihi Joshi, Xiang Ren, Swabha Swayamdipta, Rik Koncel-Kedziorski, and Tim Paek. 2025 · 2025
Closest in time.
Eliciting human preferences with language models
Belinda Z. Li, Alex Tamkin, Noah Goodman, and Jacob Andreas. 2025 · 2025
Closest in time.
The realhumaneval: Evaluating large language models’ abilities to support programmers
Hussein Mozannar, Valerie Chen, Mohammed Alsobay, Subhro Das, Sebastian Zhao, Dennis Wei, Manish Nagireddy, Prasanna Sattigeri, Ameet Talwalkar, and David Sontag. 2025 · 2025
Closest in time.
Aligning LLMs with individual preferences via interaction
Shujin Wu, Yi R. Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji. 2025 · 2025
Closest in time.