Fetching the paper…
Reading the bibliography…
Surveys have recently gained popularity as a tool to study large language models.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 1901
Earlier work this paper cites.
Survey Methodology
Groves, R., Fowler, F., Couper, M., Lepkowski, J., Singer, E., and Tourangeau, R. (2009) · 2009
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016) · 2016
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. (2018) · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. (2019) · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J. (2019) · 2019
Earlier work this paper cites.
Unqovering stereotyping biases via underspecified questions
Li, T., Khashabi, D., Khot, T., Sabharwal, A., and Srikumar, V. (2020) · 2020
Earlier work this paper cites.
Persistent anti-muslim bias in large language models
Abid, A., Farooqi, M., and Zou, J. (2021) · 2021
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. (2021) · 2021
Earlier work this paper cites.
Retiring adult: New datasets for fair machine learning
Ding, F., Hardt, M., Miller, J., and Schmidt, L. (2021) · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Eliciting bias in question answering models through ambiguity
Mao, A., Raman, N., Shu, M., Li, E., Yang, F., and Boyd-Graber, J. (2021) · 2021
Earlier work this paper cites.
The future of coding: A comparison of hand-coding and three types of computer-assisted text analysis methods
Nelson, L. K., Burk, D., Knudsen, M., and McCall, L. (2021) · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S. (2021) · 2021
Earlier work this paper cites.
CommunityLM: Probing Partisan Worldviews from Language Models
Jiang, H., Beeferman, D., Roy, B., and Roy, D. (2022) · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y. (2022) · 2022
Earlier work this paper cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. (2022) · 2022
Cited alongside, same era.
Do ais know what the most important issue is? using language models to code open-text social survey responses at scale
Mellon, J., Bailey, J., Scott, R., Breckwoldt, J., Miori, M., and Schmedeman, P. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. (2022) · 2022
Cited alongside, same era.
Using large language models to simulate multiple humans and replicate human subject studies
Aher, G. V., Arriaga, R. I., and Kalai, A. T. (2023) · 2023
Cited alongside, same era.
Out of one, many: Using language models to simulate human samples
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. (2023) · 2023
Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs
MosaicML (2023) · 2023
Closest in time.
More human than human: Measuring chatgpt political bias
Motoki, F., Pinho Neto, V., and Rodrigues, V. (2023) · 2023
Closest in time.
OpenAI (2023) · 2023
Closest in time.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukosiute, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., Jones, A., Chen, A., Mann, B., Israel, B., Seethor, B., McKinnon, C., Olah, C., Yan, D., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Khundadze, G., Kernion, J., Landis, J., Kerr, J., Mueller, J., Hyun, J., Landau, J., Ndousse, K., Goldberg, L., Lovitt, L., Lucas, M., Sellitto, M., Zhang, M., Kingsland, N., Elhage, N., Joseph, N., Mercado, N., DasSarma, N., Rausch, O., Larson, R., McCandlish, S., Johnston, S., Kravec, S., El Showk, S., Lanham, T., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Clark, J., Bowman, S. R., Askell, A., Grosse, R., Hernandez, D., Ganguli, D., Hubinger, E., Schiefer, N., and Kaplan, J. (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pythia: a suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and Van Der Wal, O. (2023) · 2023
Cited alongside, same era.
Using GPT for Market Research
Brand, J., Israeli, A., and Ngwe, D. (2023) · 2023
Cited alongside, same era.
Dolly 12b
Databricks (2023) · 2023
Cited alongside, same era.
Can AI language models replace human participants?
Dillion, D., Tandon, N., Gu, Y., and Gray, K. (2023) · 2023
Cited alongside, same era.
Do personality tests generalize to large language models?
Dorner, F., Sühr, T., Samadi, S., and Kelava, A. (2023) · 2023
Cited alongside, same era.
Towards measuring the representation of subjective global opinions in language models
Durmus, E., Nyugen, K., Liao, T. I., Schiefer, N., Askell, A., Bakhtin, A., Chen, C., Hatfield-Dodds, Z., Hernandez, D., Joseph, N., et al. (2023) · 2023
Cited alongside, same era.
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
Feng, S., Park, C. Y., Liu, Y., and Tsvetkov, Y. (2023) · 2023
Cited alongside, same era.
Rutinowski, J., Franke, S., Endendyk, J., Dormuth, I., and Pauly, M. (2023) · 2023
Closest in time.
Demonstrations of the potential of ai-based political issue polling
Sanders, N. E., Ulinich, A., and Schneier, B. (2023) · 2023
Closest in time.
Whose opinions do language models reflect?
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023) · 2023
Closest in time.
Do llms exhibit human-like response biases? a case study in survey design
Tjuatja, L., Chen, V., Wu, S. T., Talwalkar, A., and Neubig, G. (2023) · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Closest in time.
Investigating cultural alignment of large language models
AlKhamissi, B., ElNokrashy, M., Alkhamissi, M., and Diab, M. (2024) · 2024
Closest in time.
Evaluating language models as risk scores
Cruz, A. F., Hardt, M., and Mendler-Dünner, C. (2024) · 2024
Closest in time.
Training on the test task confounds evaluation and emergence
Dominguez-Olmedo, R., Dorner, F. E., and Hardt, M. (2024) · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024) · 2024
Closest in time.
Evaluating the moral beliefs encoded in llms
Scherrer, N., Shi, C., Feder, A., and Blei, D. (2024) · 2024
Closest in time.
Wang, X., Ma, B., Hu, C., Weber-Genzel, L., Röttger, P., Kreuter, F., Hovy, D., and Plank, B. (2024) · 2024
Closest in time.
Can large language models transform computational social science?
Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., and Yang, D. (2024) · 2024
Closest in time.