Fetching the paper…
Reading the bibliography…
Whether Large Language Models (LLMs) can outperform crowdsourcing on the data annotation task is attracting interest recently.
“Rank analysis of incomplete block designs: I. the method of paired comparisons,”
Ralph Allan Bradley and Milton E Terry, · 1952
Earlier work this paper cites.
“Maximum likelihood estimation of observer error-rates using the em algorithm,”
A. P. Dawid and A. M. Skene, · 1979
Earlier work this paper cites.
“Cheap and fast—but is it good?: Evaluating non-expert annotations for natural language tasks,”
R. Snow, B. O’Connor, D. Jurafsky, and A. Y. Ng, · 2008
Earlier work this paper cites.
“The multidimensional wisdom of crowds,”
Peter Welinder, Steve Branson, Serge Belongie, and Pietro Perona, · 2010
Earlier work this paper cites.
“Pairwise ranking aggregation in a crowdsourced setting,”
Xi Chen, Paul N. Bennett, Kevyn Collins-Thompson, and Eric Horvitz, · 2013
Earlier work this paper cites.
“Community-based bayesian aggregation models for crowdsourcing,”
Matteo Venanzi, John Guiver, Gabriella Kazai, Pushmeet Kohli, and Milad Shokouhi, · 2014
Earlier work this paper cites.
“A confidence-aware approach for truth discovery on long-tail data,”
Qi Li, Yaliang Li, Jing Gao, Lu Su, Bo Zhao, Murat Demirbas, Wei Fan, and Jiawei Han, · 2014
Earlier work this paper cites.
“Weather sentiment - amazon mechanical turk dataset,”
Matteo Venanzi, WTL Teacy, Alex Rogers, and Nicholas R Jennings, · 2015
Earlier work this paper cites.
“Hyper questions: Unsupervised targeting of a few experts in crowdsourcing,”
Jiyi Li, Yukino Baba, and Hisashi Kashima, · 2017
Earlier work this paper cites.
“Incorporating worker similarity for label aggregation in crowdsourcing,”
Jiyi Li, Yukino Baba, and Hisashi Kashima, · 2018
Cited alongside, same era.
“Simultaneous clustering and ranking from pairwise comparisons,”
Jiyi Li, Yukino Baba, and Hisashi Kashima, · 2018
Cited alongside, same era.
“A dataset of crowdsourced word sequences: Collections and answer aggregation for ground truth creation,”
Jiyi Li and Fumiyo Fukumoto, · 2019
Cited alongside, same era.
“Performance as a constraint: An improved wisdom of crowds using performance regularization,”
Jiyi Li, Yasushi Kawase, Yukino Baba, and Hisashi Kashima, · 2020
Cited alongside, same era.
“Rank aggregation via heterogeneous thurstone preference models,”
Tao Jin, Pan Xu, Quanquan Gu, and Farzad Farnoud, · 2020
Cited alongside, same era.
“Crowdsourced text sequence aggregation based on hybrid reliability and representation,”
“Artificial artificial artificial intelligence: Crowd workers widely use large language models for text production tasks,”
Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West, · 2023
Later among the works it cites.
“Can chatgpt reproduce human-generated labels? a study of social computing tasks,”
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson, · 2023
Later among the works it cites.
“Chatgpt outperforms crowd-workers for text-annotation tasks,”
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli, · 2023
Later among the works it cites.
“Chatgpt-4 outperforms experts and crowd workers in annotating political twitter messages with zero-shot learning,”
Petter Törnberg, · 2023
Later among the works it cites.
“Chatgpt to replace crowdsourcing of paraphrases for intent classification: Higher diversity and comparable model robustness,”
Jan Cegin, Jakub Simko, and Peter Brusilovsky, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiyi Li, · 2020
Cited alongside, same era.
“Teach me to explain: A review of datasets for explainable natural language processing,”
Sarah Wiegreffe and Ana Marasovic, · 2021
Cited alongside, same era.
“Label aggregation for crowdsourced triplet similarity comparisons,”
Jiyi Li, Lucas Ryo Endo, and Hisashi Kashima, · 2021
Cited alongside, same era.
“Context-based collective preference aggregation for prioritizing crowd opinions in social decision-making,”
Jiyi Li, · 2022
Cited alongside, same era.
“Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,”
Wei-Lin Chiang, Zhuohan Li, Zi Lin, and et al., · 2023
Later among the works it cites.
“Multiview representation learning from crowdsourced triplet comparisons,”
Xiaotian Lu, Jiyi Li, Koh Takeuchi, and Hisashi Kashima, · 2023
Later among the works it cites.
“Whose vote should count more: Optimal integration of labels from labelers of unknown expertise,”
Jacob Whitehill, Paul Ruvolo, Tingfan Wu, Jacob Bergsma, and Javier Movellan, · 2043
Closest in time.