Fetching the paper…
Reading the bibliography…
The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Subjectivity in language
Emile Benveniste. 1971 · 1971
Earlier work this paper cites.
A theory of humor
Thomas C Veatch. 1998 · 1998
Earlier work this paper cites.
Basic emotions
Paul Ekman et al. 1999 · 1999
Earlier work this paper cites.
Data augmentation using pre-trained transformer models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020 · 2003
Earlier work this paper cites.
Colbert: Using bert sentence embedding for humor detection
Issa Annamoradnejad and Gohar Zoghi. 2020 · 2004
Earlier work this paper cites.
Learning subjective language
Janyce Wiebe, Theresa Wilson, Rebecca Bruce, Matthew Bell, and Melanie Martin. 2004 · 2004
Earlier work this paper cites.
Review spam detection
Nitin Jindal and Bing Liu. 2007 · 2007
Earlier work this paper cites.
Contributions to the study of sms spam filtering: New collection and results
Tiago A. Almeida, Jose Maria Gomez Hidalgo, and Akebo Yamakami. 2011 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala. 2014 · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Semeval-2018 task 1: Affect in tweets
Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018 · 2018
Earlier work this paper cites.
Semeval-2018 task 3: Irony detection in english tweets
Cynthia Van Hee, Els Lefever, and Véronique Hoste. 2018 · 2018
Earlier work this paper cites.
FewRel 2.0: Towards more challenging few-shot relation classification
Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2019 · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Detection of abusive language: the problem of biased datasets
Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019 · 2019
Cited alongside, same era.
This dataset does not exist: training models from generated images
Victor Besnier, Himalaya Jain, Andrei Bursuc, Matthieu Cord, and Patrick Pérez. 2020 · 2020
Cited alongside, same era.
GoEmotions: A Dataset of Fine-Grained Emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Jury learning: Integrating dissenting voices into machine learning models
Mitchell L Gordon, Michelle S Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S Bernstein. 2022 · 2022
Later among the works it cites.
Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation
Nitesh Goyal, Ian D Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022 · 2022
Later among the works it cites.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Later among the works it cites.
Is synthetic data from generative models ready for image recognition?
Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue, Wenqing Zhang, Philip Torr, Song Bai, and Xiaojuan Qi. 2022 · 2022
Later among the works it cites.
Towards better detection of biased language with scarce, noisy, and biased annotations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Cited alongside, same era.
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A Smith, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Sculpting Data for ML: The first act of Machine Learning
Rishabh Misra and Jigyasa Grover. 2021 · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021 · 2021
Cited alongside, same era.
Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation
Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021 · 2021
Cited alongside, same era.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A Smith. 2021 · 2021
Cited alongside, same era.
Towards zero-label language learning
Zirui Wang, Adams Wei Yu, Orhan Firat, and Yuan Cao. 2021 · 2021
Cited alongside, same era.
Zhuoyan Li, Zhuoran Lu, and Ming Yin. 2022 · 2022
Later among the works it cites.
Generating training data with language models: Towards zero-shot language understanding
Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022 · 2022
Later among the works it cites.
Data augmentation for intent classification with off-the-shelf large language models
Gaurav Sahu, Pau Rodriguez, Issam H. Laradji, Parmida Atighehchian, David Vazquez, and Dzmitry Bahdanau. 2022 · 2022
Later among the works it cites.
Zerogen: Efficient zero-shot learning via dataset generation
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022 · 2022
Later among the works it cites.
sentence-transformers/all-minilm-l6-v2
all MiniLM-L6-v2. 2023 · 2023
Closest in time.
John Joon Young Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Closest in time.
On the creativity of large language models
Giorgio Franceschelli and Mirco Musolesi. 2023 · 2023
Closest in time.
Self-guided noise-free data generation for efficient zero-shot learning
Jiahui Gao, Renjie Pi, LIN Yong, Hang Xu, Jiacheng Ye, Zhiyong Wu, WEIZHONG ZHANG, Xiaodan Liang, Zhenguo Li, and Lingpeng Kong. 2023 · 2023
Closest in time.
Evaluating large language models in generating synthetic hci research data: A case study
Perttu Hämäläinen, Mikke Tavast, and Anton Kunnari. 2023 · 2023
Closest in time.
On the design of ai-powered code assistants for notebooks
Andrew M Mcnutt, Chenglong Wang, Robert A Deline, and Steven M. Drucker. 2023 · 2023
Closest in time.
Sarcasm detection using news headlines dataset
Rishabh Misra and Prahal Arora. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Does synthetic data generation of llms help clinical text mining?
Ruixiang Tang, Xiaotian Han, Xiaoqian Jiang, and Xia Hu. 2023 · 2023
Closest in time.
Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions
Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. 2023 · 2023
Closest in time.