Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are susceptible to persuasion, which can pose risks when models are faced with an adversarial interlocutor.
Belief, attitude, intention, and behavior: An introduction to theory and research
Martin Fishbein and Icek Ajzen. 1975 · 1975
Earlier work this paper cites.
Telling more than we can know: Verbal reports on mental processes
Richard E Nisbett and Timothy D Wilson. 1977 · 1977
Earlier work this paper cites.
Pooling of unshared information in group decision making: Biased information sampling during discussion
Garold Stasser and William Titus. 1985 · 1985
Earlier work this paper cites.
The elaboration likelihood model of persuasion
Richard E Petty and John T Cacioppo. 1986 · 1986
Earlier work this paper cites.
Evaluation may be easier than generation
Moni Naor. 1996 · 1996
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom. 2006 · 2006
Earlier work this paper cites.
Evidence for a collective intelligence factor in the performance of human groups
Anita Williams Woolley, Christopher F Chabris, Alex Pentland, Nada Hashmi, and Thomas W Malone. 2010 · 2010
Earlier work this paper cites.
Predicting and changing behavior: The reasoned action approach
Martin Fishbein and Icek Ajzen. 2011 · 2011
Earlier work this paper cites.
Sources of method bias in social science research and recommendations on how to control it
Philip M Podsakoff, Scott B MacKenzie, and Nathan P Podsakoff. 2012 · 2012
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Exploring the role of prior beliefs for argument persuasion
Esin Durmus and Claire Cardie. 2018 · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Modeling the factors of user success in online debate
Esin Durmus and Claire Cardie. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Cited alongside, same era.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec. 2020 · 2020
Cited alongside, same era.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021 · 2021
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Cited alongside, same era.
Entity-based knowledge conflicts in question answering
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2023 · 2023
Later among the works it cites.
Llama 3.1 model card
AI@Meta. 2024 · 2024
Closest in time.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al. 2024 · 2024
Closest in time.
ReConcile: Round-table conference improves reasoning via consensus among diverse LLMs
Justin Chen, Swarnadeep Saha, and Mohit Bansal. 2024 · 2024
Closest in time.
QLoRA: Efficient finetuning of quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2024 · 2024
Closest in time.
Are language models rational? the case of coherence norms and belief revision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
minicons: Enabling flexible behavioral and representational analyses of transformer language models
Kanishka Misra. 2022 · 2022
Cited alongside, same era.
Don’t generate, discriminate: A proposal for grounding language models to real-world environments
Yu Gu, Xiang Deng, and Yu Su. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Are you sure? challenging llms leads to performance drops in the flipflop experiment
Philippe Laban, Lidiya Murakhovs’ka, Caiming Xiong, and Chien-Sheng Wu. 2023 · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2023 · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Thomas Hofweber, Peter Hase, Elias Stengel-Eskin, and Mohit Bansal. 2024 · 2024
Closest in time.
Debating with more persuasive llms leads to more truthful answers
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R Bowman, Tim Rocktäschel, and Ethan Perez. 2024 · 2024
Closest in time.
Medical decision making
Harold C Sox, Michael C Higgins, Douglas K Owens, and Gillian Sanders Schmidler. 2024 · 2024
Closest in time.
Lacie: Listener-aware finetuning for confidence calibration in large language models
Elias Stengel-Eskin, Peter Hase, and Mohit Bansal. 2024 · 2024
Closest in time.
What evidence do language models find convincing?
Alexander Wan, Eric Wallace, and Dan Klein. 2024 · 2024
Closest in time.
Clasheval: Quantifying the tug-of-war between an llm’s internal prior and external evidence
Kevin Wu, Eric Wu, and James Zou. 2024 · 2024
Closest in time.
The earth is flat because…: Investigating LLMs’ belief towards misinformation via persuasive conversation
Rongwu Xu, Brian Lin, Shujian Yang, Tianqi Zhang, Weiyan Shi, Tianwei Zhang, Zhixuan Fang, Wei Xu, and Han Qiu. 2024 · 2024
Closest in time.
A survey on recent advances in llm-based multi-turn dialogue systems
Zihao Yi, Jiarui Ouyang, Yuwen Liu, Tianhao Liao, Zhe Xu, and Ying Shen. 2024 · 2024
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024 · 2024
Closest in time.