Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly used as agents that interact with users and with the world.
A behavioral model of rational choice
H. A. Simon · 1955
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases: Biases in judgments reveal some heuristics of thinking under uncertainty
A. Tversky and D. Kahneman · 1974
Earlier work this paper cites.
Mental models in cognitive science
P. N. Johnson-Laird · 1980
Earlier work this paper cites.
Probabilistic models of language processing and acquisition
N. Chater and C. D. Manning · 2006
Earlier work this paper cites.
Probabilistic models of cognition: Conceptual foundations
N. Chater, J. B. Tenenbaum, and A. Yuille · 2006
Earlier work this paper cites.
Theory-based bayesian models of inductive learning and reasoning
J. B. Tenenbaum, T. L. Griffiths, and C. Kemp · 2006
Earlier work this paper cites.
Topics in semantic association
T. L. Griffiths, M. Steyvers, and J. B. Tenenbaum · 2007
Earlier work this paper cites.
Word learning as Bayesian inference
F. Xu and J. B. Tenenbaum · 2007
Earlier work this paper cites.
Probability matching and strategy availability
D. J. Koehler and G. James · 2010
Earlier work this paper cites.
Bayesian theory of mind: Modeling joint belief-desire attribution
C. Baker, R. Saxe, and J. Tenenbaum · 2011
Earlier work this paper cites.
How to grow a mind: Statistics, structure, and abstraction
J. B. Tenenbaum, C. Kemp, T. L. Griffiths, and N. D. Goodman · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Causality in thought
S. A. Sloman and D. Lagnado · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush · 2016
Earlier work this paper cites.
Do people reason rationally about causally related events? Markov violations, weak inferences, and failures of explaining away
B. M. Rottman and R. Hastie · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
People learn other people’s preferences through inverse decision-making
A. Jern, C. G. Lucas, and C. Kemp · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Prolific.ac—A subject pool for online experiments
S. Palan and C. Schitter · 2018
Earlier work this paper cites.
Beyond Markov: Accounting for independence violations in causal reasoning
B. Rehder · 2018
Earlier work this paper cites.
The emergence of number and syntax units in LSTM language models
Y. Lakretz, G. Kruszewski, T. Desbordes, D. Hupkes, S. Dehaene, and M. Baroni · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
C. D. Manning, K. Clark, J. Hewitt, U. Khandelwal, and O. Levy · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. M. Shieber · 2020
Earlier work this paper cites.
Building and evaluating open-domain dialogue corpora with clarifying questions
M. Aliannejadi, J. Kiseleva, A. Chuklin, J. Dalton, and M. Burtsev · 2021
Cited alongside, same era.
Causal analysis of syntactic agreement mechanisms in neural language models
M. Finlayson, A. Mueller, S. Gehrmann, S. Shieber, T. Linzen, and Y. Belinkov · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al · 2021
Cited alongside, same era.
Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction
S. Ravfogel, G. Prasad, T. Linzen, and Y. Goldberg · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Perceptions of linguistic uncertainty by language models and humans
C. G. Belém, M. Kelly, M. Steyvers, S. Singh, and P. Smyth · 2024
Later among the works it cites.
Learning to maximize mutual information for chain-of-thought distillation
X. Chen, H. Huang, Y. Gao, Y. Wang, J. Zhao, and K. Ding · 2024
Later among the works it cites.
Chatbot arena: An open platform for evaluating llms by human preference
W. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. I. Jordan, J. E. Gonzalez, and I. Stoica · 2024
Later among the works it cites.
A systematic comparison of syllogistic reasoning in humans and language models
T. Eisape, M. Tessler, I. Dasgupta, F. Sha, S. Steenkiste, and T. Linzen · 2024
Later among the works it cites.
BIRD: A trustworthy bayesian inference framework for large language models
Y. Feng, B. Zhou, W. Lin, and D. Roth · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language models are zero-shot reasoners
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2022
Cited alongside, same era.
A path towards autonomous machine intelligence
Y. LeCun · 2022
Cited alongside, same era.
Inferring rewards from language in context
J. Lin, D. Fried, D. Klein, and A. Dragan · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. V. Nayak, D. Datta, J. Chang, M. T. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Févry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, and A. M. Rush · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2022
Cited alongside, same era.
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al · 2024
Later among the works it cites.
Bayesian Models of Cognition: Reverse Engineering the Mind
T. L. Griffiths, N. Chater, and J. B. Tenenbaum · 2024
Later among the works it cites.
Bayesian preference elicitation with language models
K. Handa, Y. Gal, E. Pavlick, N. Goodman, J. Andreas, A. Tamkin, and B. Z. Li · 2024
Later among the works it cites.
Impossible distillation for paraphrasing and summarization: How to make high-quality lemonade out of small, low-quality model
J. Jung, P. West, L. Jiang, F. Brahman, X. Lu, J. Fisher, T. Sorensen, and Y. Choi · 2024
Later among the works it cites.
Understanding catastrophic forgetting in language models via implicit inference
S. Kotha, J. M. Springer, and A. Raghunathan · 2024
Later among the works it cites.
Mitigating the alignment tax of RLHF
Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang, H. Dong, R. Pi, H. Zhao, N. Jiang, H. Ji, Y. Yao, and T. Zhang · 2024
Later among the works it cites.
Large language models assume people are more rational than we really are
R. Liu, J. Geng, J. Peterson, I. Sucholutsky, and T. L. Griffiths · 2024
Later among the works it cites.
What are the odds? language models are capable of probabilistic reasoning
A. Paruchuri, J. Garrison, S. Liao, J. B. Hernandez, J. Sunshine, T. Althoff, X. Liu, and D. McDuff · 2024
Later among the works it cites.
Pragmatic feature preferences: Learning reward-relevant preferences from human input
A. Peng, Y. Sun, T. Shu, and D. Abel · 2024
Later among the works it cites.
Doing experiments and revising rules with natural language and probabilistic reasoning
T. Piriyakulkij, C. Langenfeld, T. A. Le, and K. Ellis · 2024
Later among the works it cites.
On the loss of context-awareness in general instruction fine-tuning
Y. Wang, A. Bai, N. Peng, and C.-J. Hsieh · 2024
Later among the works it cites.
Unveiling the generalization power of fine-tuned large language models
H. Yang, Y. Zhang, J. Xu, H. Lu, P.-A. Heng, and W. Lam · 2024
Later among the works it cites.
Grounding language about belief in a bayesian theory-of-mind
L. Ying, T. Zhi-Xuan, L. Wong, V. Mansinghka, and J. Tenenbaum · 2024
Later among the works it cites.
Distilling system 2 into system 1
P. Yu, J. Xu, J. E. Weston, and I. Kulikov · 2024
Later among the works it cites.
Incoherent probability judgments in large language models
J.-Q. Zhu and T. Griffiths · 2024
Later among the works it cites.
Breaking the chains of independence: A bayesian uncertainty model of normative violations in human causal probabilistic reasoning
S. Chaigneau, N. Marchant, and B. Rehder · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in LLMs via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Reasoning over uncertain text by generative large language models
A. Nafar, K. B. Venable, and P. Kordjamshidi · 2025
Closest in time.
Introducing GPT-4.1 in the API, 2025
OpenAI · 2025
Closest in time.
Do LLMs recognize your preferences? evaluating personalized preference following in LLMs
S. Zhao, M. Hong, Y. Liu, D. Hazarika, and K. Lin · 2025
Closest in time.