Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown high agreement with human raters across a variety of tasks, demonstrating potential to ease the challenges of human data collection.
Mixed-initiative interaction
James E Allen, Curry I Guinn, and Eric Horvitz. 1999 · 1999
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart. 2000 · 2000
Earlier work this paper cites.
Gender differences in language use: An analysis of 14,000 text samples
Matthew L Newman, Carla J Groom, Lori D Handelman, and James W Pennebaker. 2008 · 2008
Earlier work this paper cites.
A computational approach to politeness with application to social factors
Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Dan Jurafsky, Jure Leskovec, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Measuring ideological proportions in political speeches
Yanchuan Sim, Brice D. L. Acree, Justin H. Gross, and Noah A. Smith. 2013 · 2013
Earlier work this paper cites.
An attack on science? media use, trust in scientists, and perceptions of global warming
Jay D Hmielowski, Lauren Feldman, Teresa A Myers, Anthony Leiserowitz, and Edward Maibach. 2014 · 2014
Earlier work this paper cites.
XGBoost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Language from police body camera footage shows racial disparities in officer respect
Rob Voigt, Nicholas P Camp, Vinodkumar Prabhakaran, William L Hamilton, Rebecca C Hetey, Camilla M Griffiths, David Jurgens, Dan Jurafsky, and Jennifer L Eberhardt. 2017 · 2017
Earlier work this paper cites.
The importance of calibration for estimating proportions from annotations
Dallas Card and Noah A Smith. 2018 · 2018
Earlier work this paper cites.
Uncertainty-aware generative models for inferring document class prevalence
Katherine Keith and Brendan O’Connor. 2018 · 2018
Earlier work this paper cites.
Media bias monitor: Quantifying biases of social media news outlets at large-scale
Filipe Ribeiro, Lucas Henrique, Fabricio Benevenuto, Abhijnan Chakraborty, Juhi Kulshrestha, Mahmoudreza Babaei, and Krishna Gummadi. 2018 · 2018
Earlier work this paper cites.
Auditing partisan audience bias within google search
Ronald E Robertson, Shan Jiang, Kenneth Joseph, Lisa Friedland, David Lazer, and Christo Wilson. 2018 · 2018
Earlier work this paper cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A Smith. 2019 · 2019
Earlier work this paper cites.
We can detect your bias: Predicting the political ideology of news articles
Ramy Baly, Giovanni Da San Martino, James Glass, and Preslav Nakov. 2020 · 2020
Earlier work this paper cites.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Earlier work this paper cites.
Detecting stance in media on global warming
Yiwei Luo, Dallas Card, and Dan Jurafsky. 2020 · 2020
Earlier work this paper cites.
Active learning by acquiring contrastive examples
Katerina Margatina, Giorgos Vernikos, Loïc Barrault, and Nikolaos Aletras. 2021 · 2021
Earlier work this paper cites.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Amir Feder, Katherine A Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E Roberts, et al. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022 · 2022
Earlier work this paper cites.
Taxonomy of risks posed by language models
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Marked personas: Using natural language prompts to measure stereotypes in language models
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023 · 2023
Cited alongside, same era.
Can large language models be an alternative to human evaluations?
Cheng-Han Chiang and Hung-yi Lee. 2023 · 2023
Cited alongside, same era.
When the majority is wrong: Modeling annotator disagreement for subjective tasks
Eve Fleisig, Rediet Abebe, and Dan Klein. 2023 · 2023
Cited alongside, same era.
ChatGPT outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023 · 2023
Cited alongside, same era.
LLMaAA: Making large language models as active annotators
Ruoyu Zhang, Yanzeng Li, Yongliang Ma, Ming Zhou, and Lei Zou. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023 · 2023
Later among the works it cites.
Normbank: A knowledge bank of situational social norms
Caleb Ziems, Jane Dwivedi-Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Prompt design matters for computational social science tasks but in unpredictable ways
Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, and Libby Hemphill. 2024 · 2024
Closest in time.
Can generative AI improve social science?
Christopher A Bail. 2024 · 2024
Closest in time.
Linguistic calibration of long-form generations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Culturally aware natural language inference
Jing Huang and Diyi Yang. 2023 · 2023
Cited alongside, same era.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023 · 2023
Cited alongside, same era.
Auditing the AI auditors: A framework for evaluating fairness and bias in high stakes AI predictive models
Richard N Landers and Tara S Behrend. 2023 · 2023
Cited alongside, same era.
The dimensions of data labor: A road map for researchers, activists, and policymakers to empower data producers
Hanlin Li, Nicholas Vincent, Stevie Chancellor, and Brent Hecht. 2023a · 2023
Cited alongside, same era.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023b · 2023
Cited alongside, same era.
Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation
Minzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan, Nancy Chen, Zhengyuan Liu, and Diyi Yang. 2023c · 2023
Cited alongside, same era.
G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
Neil Band, Xuechen Li, Tengyu Ma, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
Autoeval done right: Using synthetic data for model evaluation
Pierre Boyeau, Anastasios N Angelopoulos, Nir Yosef, Jitendra Malik, and Michael I Jordan. 2024 · 2024
Closest in time.
Gen-AI: Artificial intelligence and the future of work
Mauro Cazzaniga, Ms Florence Jaumotte, Longji Li, Mr Giovanni Melina, Augustus J Panton, Carlo Pizzinelli, Emma J Rockall, and Ms Marina Mendes Tavares. 2024 · 2024
Closest in time.
Prediction-powered ranking of large language models
Ivi Chatzi, Eleni Straitouri, Suhas Thejaswi, and Manuel Gomez Rodriguez. 2024 · 2024
Closest in time.
Using imperfect surrogates for downstream inference: Design-based supervised learning for social science applications of large language models
Naoki Egami, Musashi Hinck, Brandon Stewart, and Hanying Wei. 2024 · 2024
Closest in time.
A survey of confidence estimation and calibration in large language models
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2024 · 2024
Closest in time.
Detecting and preventing hallucinations in large vision language models
Anisha Gunjal, Jihan Yin, and Erhan Bas. 2024 · 2024
Closest in time.
Llm-rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts
Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, and Chris Kedzie. 2024 · 2024
Closest in time.
MEGAnno+: A human-LLM collaborative annotation system
Hannah Kim, Kushan Mitra, Rafael Li Chen, Sajjadur Rahman, and Dan Zhang. 2024 · 2024
Closest in time.
Concept induction: Analyzing unstructured text with high-level concepts using lloom
Michelle S Lam, Janice Teoh, James A Landay, Jeffrey Heer, and Michael S Bernstein. 2024 · 2024
Closest in time.
TopicGPT: A prompt-based topic modeling framework
Chau Pham, Alexander Hoyle, Simeng Sun, Philip Resnik, and Mohit Iyyer. 2024 · 2024
Closest in time.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2024 · 2024
Closest in time.
Ghostbuster: Detecting text ghostwritten by large language models
Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2024 · 2024
Closest in time.
Relying on the unreliable: The impact of language models’ reluctance to express uncertainty
Kaitlyn Zhou, Jena D Hwang, Xiang Ren, and Maarten Sap. 2024 · 2024
Closest in time.
Can large language models transform computational social science?
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024 · 2024
Closest in time.
Active statistical inference
Tijana Zrnic and Emmanuel J Candès. 2024 · 2024
Closest in time.