Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are regularly being used to label data across many domains and for myriad tasks.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
Few-shot text generation with pattern-exploiting training
Timo Schick and Hinrich Schütze. 2020 · 2012
Earlier work this paper cites.
SemEval-2016 task 6: Detecting stance in tweets
Saif Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry. 2016 · 2016
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019 · 2019
Earlier work this paper cites.
Jigsaw unintended bias in toxicity classification
cjadams, Daniel Borkan, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and nithum. 2019 · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Neural Network Acceptability Judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
iSarcasm: A dataset of intended sarcasm
Silviu Oprea and Walid Magdy. 2020 · 2020
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2020 · 2020
Cited alongside, same era.
Warp: Word-level adversarial reprogramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Cited alongside, same era.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner. 2021 · 2021
Cited alongside, same era.
Colbert: Using bert sentence embedding in parallel neural networks for computational humor
Chatgpt: Jack of all trades, master of none
Jan Kocoń, Igor Cichecki, Oliwier Kaszyca, Mateusz Kochanek, Dominika Szydło, Joanna Baran, Julita Bielaniewicz, Marcin Gruza, Arkadiusz Janz, Kamil Kanclerz, Anna Kocoń, Bartłomiej Koptyra, Wiktoria Mieleszczenko-Kowszewicz, Piotr Miłkowski, Marcin Oleksy, Maciej Piasecki, Łukasz Radliński, Konrad Wojtasik, Stanisław Woźniak, and Przemysław Kazienko. 2023 · 2023
Later among the works it cites.
Making large language models better data creators
Dong-Ho Lee, Jay Pujara, Mohit Sewak, Ryen White, and Sujay Jauhar. 2023 · 2023
Later among the works it cites.
The hitchhiker’s guide to program analysis: A journey with large language models
Haonan Li, Yu Hao, Yizhuo Zhai, and Zhiyun Qian. 2023 · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Issa Annamoradnejad and Gohar Zoghi. 2022 · 2022
Cited alongside, same era.
Quantifying social biases using templates is unreliable
Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022 · 2022
Cited alongside, same era.
Principled instructions are all you need for questioning llama-1/2, gpt-3.5/4
Sondos Mahmoud Bsharat, Aidar Myrzakhan, and Zhiqiang Shen. 2023 · 2023
Cited alongside, same era.
Are large language model-based evaluators the solution to scaling up multilingual evaluation?
Rishav Hada, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2023 · 2023
Cited alongside, same era.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023 · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Later among the works it cites.
Can chatgpt reproduce human-generated labels? a study of social computing tasks
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. 2023 · 2023
Later among the works it cites.
Dr chatgpt, tell me what i want to hear: How prompt knowledge impacts health answer correctness
Guido Zuccon and Bevan Koopman. 2023 · 2023
Later among the works it cites.