Fetching the paper…
Reading the bibliography…
Large language models (LLMs) can be used to generate text data for training and evaluating other models.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Cueflik: Interactive concept learning in image search
James Fogarty, Desney Tan, Ashish Kapoor, and Simon Winder. 2008 · 2008
Earlier work this paper cites.
Overview based example selection in end user interactive concept learning
Saleema Amershi, James Fogarty, Ashish Kapoor, and Desney Tan. 2009 · 2009
Earlier work this paper cites.
Ensemblematrix: Interactive visualization to support machine learning with multiple classifiers
Justin Talbot, Bongshin Lee, Ashish Kapoor, and Desney S. Tan. 2009 · 2009
Earlier work this paper cites.
Visual recognition with humans in the loop
Steve Branson, Catherine Wah, Florian Schroff, Boris Babenko, Peter Welinder, Pietro Perona, and Serge Belongie. 2010 · 2010
Earlier work this paper cites.
Interactive optimization for steering machine classification
Ashish Kapoor, Bongshin Lee, Desney Tan, and Eric Horvitz. 2010 · 2010
Earlier work this paper cites.
Regroup: Interactive machine learning for on-demand group creation in social networks
Saleema Amershi, James Fogarty, and Daniel Weld. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Flock: Hybrid crowd-machine learning classifiers
Justin Cheng and Michael S. Bernstein. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Squinky! a corpus of sentence-level formality, informativeness, and implicature
Shibamouli Lahiri. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Stop clickbait: Detecting and preventing clickbaits in online news media
Abhijnan Chakraborty, Bhargavi Paranjape, Sourya Kakarla, and Niloy Ganguly. 2016 · 2016
Earlier work this paper cites.
Deep Learning
Ian J. Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Earlier work this paper cites.
PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts
Franck Dernoncourt and Ji Young Lee. 2017 · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
Earlier work this paper cites.
Sequence-to-sequence data augmentation for dialogue language understanding
Yutai Hou, Yijia Liu, Wanxiang Che, and Ting Liu. 2018 · 2018
Earlier work this paper cites.
CARER: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018 · 2018
Cited alongside, same era.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A Efros. 2018 · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert. 2019 · 2019
Cited alongside, same era.
Anchorviz: Facilitating semantic data exploration and concept discovery for interactive machine learning
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Cg-bert: Conditional text generation with bert for generalized few-shot intent detection
Congying Xia, Chenwei Zhang, Hoang Nguyen, Jiawei Zhang, and Philip Yu. 2020 · 2020
Later among the works it cites.
Discovering and validating ai errors with crowdsourced failure reports
Ángel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, and Adam Perer. 2021 · 2021
Later among the works it cites.
Benchmarking Natural Language Understanding Services for Building Conversational Agents , pages 165–183. Springer Singapore, Singapore
Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2021 · 2021
Later among the works it cites.
Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jina Suh, Soroush Ghorashi, Gonzalo Ramos, Nan-Chen Chen, Steven Drucker, Johan Verwey, and Patrice Simard. 2019 · 2019
Cited alongside, same era.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
EDA: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou. 2019 · 2019
Cited alongside, same era.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2019 · 2019
Cited alongside, same era.
Data augmentation for spoken language understanding via joint variational generation
Kang Min Yoo, Youhyun Shin, and Sang-goo Lee. 2019 · 2019
Cited alongside, same era.
MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification
Jiaao Chen, Zichao Yang, and Diyi Yang. 2020 · 2020
Cited alongside, same era.
Sequence-level mixed sample data augmentation
Demi Guo, Yoon Kim, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021 · 2021
Later among the works it cites.
GPT3Mix: Leveraging large-scale language models for text augmentation
Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woomyoung Park. 2021 · 2021
Later among the works it cites.
Synthbio: A case study in faster curation of text datasets
Ann Yuan, Daphne Ippolito, Vitaly Nikolaev, Chris Callison-Burch, Andy Coenen, and Sebastian Gehrmann. 2021 · 2021
Later among the works it cites.
Dataset distillation by matching training trajectories
George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros, and Jun-Yan Zhu. 2022 · 2022
Later among the works it cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022 · 2022
Later among the works it cites.
Trade-offs in sampling and search for early-stage interactive text classification
Zachary Levonian, Chia-Jung Lee, Vanessa Murdock, and F. Maxwell Harper. 2022 · 2022
Later among the works it cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Adaptive testing and debugging of NLP models
Marco Tulio Ribeiro and Scott Lundberg. 2022 · 2022
Later among the works it cites.
Data augmentation for intent classification with off-the-shelf large language models
Gaurav Sahu, Pau Rodriguez, Issam Laradji, Parmida Atighehchian, David Vazquez, and Dzmitry Bahdanau. 2022 · 2022
Later among the works it cites.
Isea: An interactive pipeline for semantic error analysis of nlp models
Jun Yuan, Jesse Vig, and Nazneen Rajani. 2022 · 2022
Later among the works it cites.
TreeMix: Compositional constituency-based data augmentation for natural language understanding
Le Zhang, Zichao Yang, and Diyi Yang. 2022 · 2022
Later among the works it cites.
FlipDA: Effective and robust data augmentation for few-shot learning
Jing Zhou, Yanan Zheng, Jie Tang, Li Jian, and Zhilin Yang. 2022 · 2022
Later among the works it cites.
Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future
Jan-Christoph Klie, Bonnie Webber, and Iryna Gurevych. 2023 · 2023
Closest in time.