Fetching the paper…
Reading the bibliography…
We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior.
Vox populi
Galton, F · 1907
Earlier work this paper cites.
Computing machinery and intelligence
Turing, A. M · 1950
Earlier work this paper cites.
Behavioral study of obedience
Milgram, S · 1963
Earlier work this paper cites.
An experimental analysis of ultimatum bargaining
Güth, W., Schmittberger, R., and Schwarze, B · 1982
Earlier work this paper cites.
On not being led up the garden path: the use of context by the psychological syntax processor , pp. 320–358
Crain, S. and Steedman, M · 1985
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn treebank
Marcus, M. P., Santorini, B., and Marcinkiewicz, M. A · 1993
Earlier work this paper cites.
Thematic roles assigned along the garden path linger
Christianson, K., Hollingworth, A., Halliwell, J. F., and Ferreira, F · 2001
Earlier work this paper cites.
Chivalry and Solidarity in Ultimatum Games
Eckel, C. C. and Grossman, P. J · 2001
Earlier work this paper cites.
The Wisdom of Crowds
Surowiecki, J · 2004
Earlier work this paper cites.
The Difference: How the Power of Diversity Creates Better Groups, Firms, Schools, and Societies
Page, S · 2007
Earlier work this paper cites.
Lingering misinterpretations in garden-path sentences: evidence from a paraphrasing task
Patson, N. D., Darowski, E. S., Moon, N., and Ferreira, F · 2009
Earlier work this paper cites.
Tutorial on agent-based modelling and simulation
Macal, C. and North, M · 2010
Earlier work this paper cites.
Wikipedia:Systemic bias — Wikipedia, the free encyclopedia
Wikipedia · 2010
Earlier work this paper cites.
Social influence and the collective dynamics of opinion formation
Moussaïd, M., Kämmer, J. E., Analytis, P. P., and Neth, H · 2013
Earlier work this paper cites.
Chapter 2 - experimental economics and experimental game theory
Houser, D. and McCabe, K · 2014
Earlier work this paper cites.
Suicide risk assessment and intervention in people with mental illness
Bolton, J. M., Gunnell, D., and Turecki, G · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Bolukbasi, T., Chang, K.-W., Zou, J. Y., Saligrama, V., and Kalai, A. T · 2016
Earlier work this paper cites.
Extending legal protection to social robots: The effects of anthropomorphism, empathy, and violent behavior towards robotic objects
Darling, K · 2016
Earlier work this paper cites.
Can machines think? a report on turing test experiments at the royal society
Warwick, K. and Shah, H · 2016
Cited alongside, same era.
White, man, and highly followed: Gender and race inequalities in twitter
Messias, J., Vikatos, P., and Benevenuto, F · 2017
Cited alongside, same era.
Chapter 12 - social cognition: Reasoning with others
Krawczyk, D. C · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Climbing towards nlu: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A · 2020
Cited alongside, same era.
Language (technology) is power: A critical survey of “bias” in nlp
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2020
Language models show human-like content effects on reasoning
Dasgupta, I., Lampinen, A. K., Chan, S. C., Creswell, A., Kumaran, D., McClelland, J. L., and Hill, F · 2022
Closest in time.
Machine intuition: Uncovering human-like intuitive decision-making in gpt-3.5
Hagendorff, T., Fabi, S., and Kosinski, M · 2022
Closest in time.
Mpi: Evaluating and inducing personality in pre-trained language models
Jiang, G., Xu, M., Zhu, S.-C., Han, W., Zhang, C., and Zhu, Y · 2022
Closest in time.
Capturing failures of large language models via human cognitive biases
Jones, E. and Steinhardt, J · 2022
Closest in time.
Ai personification: Estimating the personality of language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Cited alongside, same era.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Delphi: Towards machine ethics and norms
Jiang, L., Hwang, J. D., Bhagavatula, C., Bras, R. L., Forbes, M., Borchardt, J., Liang, J., Etzioni, O., Sap, M., and Choi, Y · 2021
Cited alongside, same era.
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2021
Cited alongside, same era.
Karra, S. R., Nguyen, S., and Tulabandhula, T · 2022
Closest in time.
Large language models are zero-shot reasoners, 2022
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Closest in time.
Holistic evaluation of language models, 2022
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2022
Closest in time.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Closest in time.
Social simulacra: Creating populated prototypes for social computing systems
Park, J. S., Popowski, L., Cai, C., Morris, M. R., Liang, P., and Bernstein, M. S · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2022
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., and (422-others) · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models, 2022
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2022
Closest in time.
Out of one, many: Using language models to simulate human samples
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D · 2023
Closest in time.
Using cognitive psychology to understand gpt-3
Binz, M. and Schulz, E · 2023
Closest in time.
Language models and cognitive automation for economic research
Korinek, A · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
Kosinski, M · 2023
Closest in time.
GPT-4 Technical Report, March 2023
OpenAI · 2023
Closest in time.
Large language models fail on trivial alterations to theory-of-mind tasks
Ullman, T · 2023
Closest in time.