Fetching the paper…
Reading the bibliography…
Value trade-offs are an integral part of human decision-making and language use, however, current tools for interpreting such dynamic and multi-faceted notions of values in language models are limited.
Convention: A Philosophical Study
David K. Lewis · 1969
Earlier work this paper cites.
Logic and conversation
H. Paul Grice · 1975
Earlier work this paper cites.
Ascribing mental qualities to machines
John McCarthy · 1979
Earlier work this paper cites.
Choice and consequence
Thomas C Schelling et al · 1984
Earlier work this paper cites.
Society of mind
Marvin Minsky · 1986
Earlier work this paper cites.
Intentional systems
Daniel Dennett · 1987
Earlier work this paper cites.
Consciousness Explained
Daniel Dennett · 1991
Earlier work this paper cites.
Breakdown of will
George Ainslie · 2001
Earlier work this paper cites.
Help or hinder: Bayesian models of social goal inference
Tomer Ullman, Chris Baker, Owen Macindoe, Owain Evans, Noah Goodman, and Joshua Tenenbaum · 2009
Earlier work this paper cites.
Predicting pragmatic reasoning in language games
Michael C. Frank and Noah D. Goodman · 2012
Earlier work this paper cites.
The no-u-turn sampler: adaptively setting path lengths in hamiltonian monte carlo
Matthew D Hoffman, Andrew Gelman, et al · 2014
Earlier work this paper cites.
Reasoning about social choices and social relationships
Alan Jern and Charles Kemp · 2014
Earlier work this paper cites.
Nonliteral understanding of number words
Justine T Kao, Jean Y Wu, Leon Bergen, and Noah D Goodman · 2014
Earlier work this paper cites.
Concrete problems in ai safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Pragmatic language interpretation as probabilistic inference
Noah D Goodman and Michael C Frank · 2016
Earlier work this paper cites.
The naïve utility calculus: Computational principles underlying commonsense psychology: (trends in cognitive sciences 20, 589–604; july 19, 2016)
Julian Jara-Ettinger, Hyowon Gweon, Laura E. Schulz, and Joshua B. Tenenbaum · 2016
Earlier work this paper cites.
Stan: A probabilistic programming language
Bob Carpenter, Andrew Gelman, Matthew D Hoffman, Daniel Lee, Ben Goodrich, Michael Betancourt, Marcus Brubaker, Jiqiang Guo, Peter Li, and Allen Riddell · 2017
Earlier work this paper cites.
Children change their answers in response to neutral follow‐up questions by a knowledgeable asker
Elizabeth Baraff Bonawitz, Patrick Shafto, Yue Yu, Aaron Gonzalez, and Sophie Bridgers · 2019
Earlier work this paper cites.
Theory of mind as inverse reinforcement learning
Julian Jara-Ettinger · 2019
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Polite speech emerges from competing social goals
Erica J Yoon, Michael Henry Tessler, Noah D Goodman, and Michael C Frank · 2020
Earlier work this paper cites.
Bayesian data analysis third edition (with errors fixed as of 6 april 2021)
Andrew Gelman, John B Carlin, Hal S Stern, David B Dunson, Aki Vehtari, and Donald B Rubin · 2021
Earlier work this paper cites.
A pragmatic account of the weak evidence effect
Samuel A Barnett, Thomas L Griffiths, and Robert D Hawkins · 2022
Earlier work this paper cites.
Mysteries of mode collapse
janus · 2022
Earlier work this paper cites.
Reward (mis)design for autonomous driving
W. Bradley Knox, Alessandro Allievi, Holger Banzhaf, Felix Schmitt, and Peter Stone · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Adopted utility calculus: Origins of a concept of social affiliation
Lindsey J Powell · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
How to handle the truth: A model of politeness as strategic truth-stretching
Fausto Carcassi and Michael Franke · 2023
Earlier work this paper cites.
Large language models behave (almost) as rational speech actors: Insights from metaphor understanding
Gaia Carenini, Louis Bodot, Luca Bischetti, Walter Schaeken, and Valentina Bambini · 2023
Earlier work this paper cites.
Identifying social partners through indirect prosociality: A computational account
Isaac Davis, Ryan Carlson, Yarrow Dunham, and Julian Jara-Ettinger · 2023
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Cited alongside, same era.
Prompting is not a substitute for probability measurements in large language models
Jennifer Hu and Roger Levy · 2023
Cited alongside, same era.
Language models are bounded pragmatic speakers
Khanh Xuan Nguyen · 2023
Cited alongside, same era.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn · 2023
Cited alongside, same era.
Reconciling truthfulness and relevance as epistemic and decision-theoretic utility
Theodore R. Sumers, Mark K. Ho, Thomas L. Griffiths, and Robert D. Hawkins · 2023
Cited alongside, same era.
Preference learning algorithms do not learn preference rankings
Angelica Chen, Sadhika Malladi, Lily Zhang, Xinyi Chen, Qiuyi Richard Zhang, Rajesh Ranganath, and Kyunghyun Cho · 2024
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Beyond preferences in ai alignment
Tan Zhi-Xuan, Micah Carroll, Matija Franklin, and Hal Ashton · 2024
Later among the works it cites.
Social sycophancy: A broader understanding of llm sycophancy, 2025
Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe, Lujain Ibrahim, and Dan Jurafsky · 2025
Closest in time.
Syceval: Evaluating llm sycophancy, 2025
Aaron Fanous, Jacob Goldberg, Ank A. Agarwal, Joanna Lin, Anson Zhou, Roxana Daneshjou, and Sanmi Koyejo · 2025
Closest in time.
Econevals: Benchmarks and litmus tests for llm agents in unknown environments, 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Kto: Model alignment as prospect theoretic optimization
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Cited alongside, same era.
Orpo: Monolithic preference optimization without reference model
Jiwoo Hong, Noah Lee, and James Thorne · 2024
Cited alongside, same era.
Openrlhf: An easy-to-use, scalable and high-performance rlhf framework
Jian Hu, Xibin Wu, Zilin Zhu, Xianyu, Weixun Wang, Dehao Zhang, and Yu Cao · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
What makes and breaks safety fine-tuning? a mechanistic study, 2024
Samyak Jain, Ekdeep Singh Lubana, Kemal Oksuz, Tom Joy, Philip H. S. Torr, Amartya Sanyal, and Puneet K. Dokania · 2024
Cited alongside, same era.
Sara Fish, Julia Shephard, Minkai Li, Ran I. Shorrer, and Yannai A. Gonczarowski · 2025
Closest in time.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile, and Noah D Goodman · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
We can’t understand ai using our existing vocabulary
John Hewitt, Robert Geirhos, and Been Kim · 2025
Closest in time.
Safety tax: Safety alignment makes your large reasoning models less reasonable
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu · 2025
Closest in time.
Itay Itzhak, Yonatan Belinkov, and Gabriel Stanovsky · 2025
Closest in time.
Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran · 2025
Closest in time.
Jared Joselowitz, Ritam Majumdar, Arjun Jagota, Matthieu Bou, Nyal Patel, Satyapriya Krishna, and Sonali Parbhoo · 2025
Closest in time.
Beyond interpretability: Developing a language to shape our relationships with ai
Been Kim · 2025
Closest in time.
On the biology of a large language model
Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, David Abrahams, Shan Carter, Basil Hosmer, Jonathan Marcus, Michael Sklar, Adly Templeton, Trenton Bricken, Callum McDougall, Hoagy Cunningham, Thomas Henighan, Adam Jermyn, Andy Jones, Andrew Persic, Zhenyi Qi, T. Ben Thompson, Sam Zimmerman, Kelley Rivoire, Thomas Conerly, Chris Olah, and Joshua Batson · 2025
Closest in time.
Truth decay: Quantifying multi-turn sycophancy in language models, 2025
Joshua Liu, Aarav Jain, Soham Takuri, Srihan Vege, Aslihan Akalin, Kevin Zhu, Sean O’Brien, and Vasu Sharma · 2025
Closest in time.
Auditing language models for hidden objectives, 2025
Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra-Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R. Bowman, Shan Carter, Brian Chen, Hoagy Cunningham, Carson Denison, Florian Dietz, Satvik Golechha, Akbir Khan, Jan Kirchner, Jan Leike, Austin Meek, Kei Nishimura-Gasparian, Euan Ong, Christopher Olah, Adam Pearce, Fabien Roger, Jeanne Salle, Andy Shih, Meg Tong, Drake Thomas, Kelley Rivoire, Adam Jermyn, Monte MacDiarmid, Tom Henighan, and Evan Hubinger · 2025
Closest in time.
One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity
Sonia Krishna Murthy, Tomer Ullman, and Jennifer Hu · 2025
Closest in time.
Sycophancy in gpt-4o: What happened and what we’re doing about it
OpenAI · 2025
Closest in time.
On the same wavelength? evaluating pragmatic reasoning in language models across broad concepts
Linlu Qiu, Cedegao E. Zhang, Joshua B. Tenenbaum, Yoon Kim, and Roger P. Levy · 2025
Closest in time.
Bridging the human-ai knowledge gap through concept discovery and transfer in alphazero
Lars Schut, Nenad Tomašev, Timothy McGrath, Demis Hassabis, Ulrich Paquet, and Been Kim · 2025
Closest in time.
Yunyi Shen and Tamara Broderick · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al · 2025
Closest in time.
Integrating neural and symbolic components in a model of pragmatic question-answering
Polina Tsvilodub, Robert D. Hawkins, and Michael Franke · 2025
Closest in time.
Base Models Beat Aligned Models at Randomness and Creativity, 2025
Peter West and Christopher Potts · 2025
Closest in time.
Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad, Stephan Günnemann, and Johannes Gasteiger · 2025
Closest in time.
Weihao Zeng, Yuzhen Huang, Qian Liu, Wei Liu, Keqing He, Zejun Ma, and Junxian He · 2025
Closest in time.
Comparing human and llm politeness strategies in free production, 2025
Haoran Zhao and Robert D. Hawkins · 2025
Closest in time.
Echo chamber: Rl post-training amplifies behaviors learned in pretraining, 2025
Rosie Zhao, Alexandru Meterez, Sham Kakade, Cengiz Pehlevan, Samy Jelassi, and Eran Malach · 2025
Closest in time.
The hidden risks of large reasoning models: A safety assessment of r1
Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Shreedhar Jangam, Jayanth Srinivasa, Gaowen Liu, Dawn Song, and Xin Eric Wang · 2025
Closest in time.
Representation engineering: A top-down approach to ai transparency, 2025
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks · 2025
Closest in time.
Reward models inherit value biases from pretraining, 2026
Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang, Vincent Adam, Hannah Rose Kirk, Christopher Summerfield, and Tsvetomira Dumbalska · 2026
Closest in time.