Fetching the paper…
Reading the bibliography…
Language models (LMs) are increasingly being used in open-ended contexts, where the opinions reflected by LMs in response to subjective queries can have a profound impact, both on user satisfaction, as well as shaping the views of society at large.
Studies in public opinion: Attitudes, nonattitudes, measurement error, and change
Saris, W. and Sniderman, P · 2004
Earlier work this paper cites.
The weirdest people in the world?
Henrich, J., Heine, S., and Norenzayan, A · 2010
Earlier work this paper cites.
Subjective natural language problems: Motivations, applications, characterizations, and implications
Alm, C · 2011
Earlier work this paper cites.
Measuring public opinion with surveys
Berinsky, A · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
De-Arteaga, M., Romanov, A., Wallach, H., Chayes, J., Borgs, C., Chouldechova, A., Geyik, S., Kenthapadi, K., and Kalai, A. T · 2019
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Pavlick, E. and Kwiatkowski, T · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
How can we know what language models know?
Jiang, F., Xu, F., Araki, J., and Neubig, G · 2020
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., en, B. M., DasSarma, N., et al · 2021
Earlier work this paper cites.
BOLD: Dataset and metrics for measuring biases in open-ended language generation
Dhamala, J., Sun, T., Kumar, V., Krishna, S., Pruksachatkun, Y., Chang, K., and Gupta, R · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Earlier work this paper cites.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Gordon, M., Zhou, K., Patel, K., Hashimoto, T. B., and Bernstein, M · 2021
Cited alongside, same era.
Jurassic-1: Technical details and evaluation
Lieber, O., Sharir, O., Lenz, B., and Shoham, Y · 2021
Cited alongside, same era.
Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes
Lourie, N., Bras, R. L., and Choi, Y · 2021
Cited alongside, same era.
Turingbench: A benchmark environment for turing test in the age of neural text generation
Uchendu, A., Ma, Z., Le, T., Zhang, R., and Lee, D · 2021
Cited alongside, same era.
Bot-adversarial dialogue for safe conversational agents
Xu, J., Ju, D., Li, M., Boureau, Y., Weston, J., and Dinan, E · 2021
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Later among the works it cites.
Jury learning: Integrating dissenting voices into machine learning models
Gordon, M., Lam, M., Park, J., Patel, K., Hancock, J., Hashimoto, T. B., and Bernstein, M · 2022
Later among the works it cites.
Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation
Goyal, N., Kivlichan, I., Rosen, R., and Vasserman, L · 2022
Later among the works it cites.
Communitylm: Probing partisan worldviews from language models
Jiang, H., Beeferman, D., Roy, B., and Roy, D · 2022
Later among the works it cites.
Surfacing racial stereotypes through identity portrayal
Kambhatla, G., Stewart, I., and Mihalcea, R · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yan, T · 2021
Cited alongside, same era.
Using large language models to simulate multiple humans
Aher, G., Arriaga, R., and Kalai, A · 2022
Cited alongside, same era.
Jurassic-1 Instruct [beta]
AI21Labs · 2022
Cited alongside, same era.
Out of one, many: Using language models to simulate human samples
Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J., Rytting, C., and Wingate, D · 2022
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Cited alongside, same era.
Using cognitive psychology to understand gpt-3
Binz, M. and Schulz, E · 2022
Cited alongside, same era.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Davani, D., Díaz, M., and Prabhakaran, V · 2022
Cited alongside, same era.
Later among the works it cites.
Ai personification: Estimating the personality of language models
Karra, S., Nguyen, S., and Tulabandhula, T · 2022
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Social simulacra: Creating populated prototypes for social computing systems
Park, J. S., Popowski, L., Cai, C., Morris, M. R., Liang, P., and Bernstein, M. S · 2022
Later among the works it cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Sap, M., Swayamdipta, S., Vianna, L., Zhou, X., Choi, Y., and Smith, N. A · 2022
Later among the works it cites.
Moral mimicry: Large language models produce moral rationalizations tailored to political identity
Simmons, G · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A., Abid, A., Fisch, A., Brown, A., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Hartmann, J., Schwenzow, J., and Witte, M · 2023
Closest in time.