Fetching the paper…
Reading the bibliography…
The latest generative large language models (LLMs) have found their application in data augmentation tasks, where small numbers of text samples are LLM-paraphrased and then used to fine-tune downstream models.
ChatGPT to replace crowdsourcing of paraphrases for intent classification: Higher diversity and comparable model robustness
Jan Cegin, Jakub Simko, and Peter Brusilovsky. 2023 · 1905
Earlier work this paper cites.
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. 2019 · 1910
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang. 1995 · 1995
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Effective quality assurance for data labels through crowdsourcing and domain expert collaboration
Lee Wei, Hsuan Huang Chi, Wei Chang Chien, Kuang Daniel Wu Ming, Ta Chuang Kun, An Yang Po, and Cheng Hsieh Chu. 2018 · 2018
Earlier work this paper cites.
Outlier detection for improved data quality and diversity in dialog systems
Stefan Larson, Anish Mahendran, Andrew Lee, Jonathan K. Kummerfeld, Parker Hill, Michael A. Laurenzano, Johann Hauswald, Lingjia Tang, and Jason Mars. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Cross-lingual transfer learning for multilingual task oriented dialog
Sebastian Schuster, Sonal Gupta, Rushin Shah, and Mike Lewis. 2019 · 2019
Earlier work this paper cites.
A semantically consistent and syntactically variational encoder-decoder framework for paraphrase generation
Wenqing Chen, Jidong Tian, Liqiang Xiao, Hao He, and Yaohui Jin. 2020 · 2020
Earlier work this paper cites.
Neural syntactic preordering for controlled paraphrase generation
Tanya Goyal and Greg Durrett. 2020 · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
Kalpesh Krishna, John Wieting, and Mohit Iyyer. 2020 · 2020
Earlier work this paper cites.
Iterative feature mining for constraint-based data collection to increase data diversity and model robustness
Stefan Larson, Anthony Zheng, Anish Mahendran, Rishi Tekriwal, Adrian Cheung, Eric Guldan, Kevin Leach, and Jonathan K. Kummerfeld. 2020 · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. 2020 · 2020
Cited alongside, same era.
Paraphrase generation as zero-shot multilingual translation: Disentangling semantic similarity from lexical and syntactic diversity
Brian Thompson and Matt Post. 2020 · 2020
Cited alongside, same era.
Dynamic word recommendation to obtain diverse crowdsourced paraphrases of user utterances
Jianing Zhou and Suma Bhat. 2020 · 2020
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Cited alongside, same era.
Dale: Generative data augmentation for low-resource legal nlp
Sreyan Ghosh, Chandra Kiran Evuru, Sonal Kumar, S Ramaneswaran, S Sakshi, Utkarsh Tyagi, and Dinesh Manocha. 2023 · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Later among the works it cites.
MEAL: Stable and active learning for few-shot prompting
Abdullatif Köksal, Timo Schick, and Hinrich Schuetze. 2023 · 2023
Later among the works it cites.
Platypus: Quick, cheap, and powerful refinement of llms
Ariel N. Lee, Cole J. Hunter, and Nataniel Ruiz. 2023 · 2023
Later among the works it cites.
Task contamination: Language models may not be few-shot anymore
Changmao Li and Jeffrey Flanigan. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Directed diversity: Leveraging language embedding distances for collective creativity in crowd ideation
Samuel Rhys Cox, Yunlong Wang, Ashraf Abdul, Christian von der Weth, and Brian Y. Lim. 2021 · 2021
Cited alongside, same era.
Novelty controlled paraphrase generation with retrieval augmented conditional prompt tuning
Jishnu Ray Chowdhury, Yong Zhuang, and Shuyi Wang. 2022 · 2022
Cited alongside, same era.
An investigation of the (in)effectiveness of counterfactually augmented data
Nitish Joshi and He He. 2022 · 2022
Cited alongside, same era.
Data curation alone can stabilize in-context learning
Ting-Yun Chang and Robin Jia. 2023 · 2023
Cited alongside, same era.
Prompting a large language model to generate diverse motivational messages: A comparison with human-written messages
Samuel Rhys Cox, Ashraf Abdul, and Wei Tsang Ooi. 2023 · 2023
Cited alongside, same era.
Auggpt: Leveraging chatgpt for text data augmentation
Haixing Dai, Zhengliang Liu, Wenxiong Liao, Xiaoke Huang, Yihan Cao, Zihao Wu, Lin Zhao, Shaochen Xu, Wei Liu, Ninghao Liu, Sheng Li, Dajiang Zhu, Hongmin Cai, Lichao Sun, Quanzheng Li, Dinggang Shen, Tianming Liu, and Xiang Li. 2023 · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Finding support examples for in-context learning
Xiaonan Li and Xipeng Qiu. 2023 · 2023
Later among the works it cites.
Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation
Marius Mosbach, Tiago Pimentel, Shauli Ravfogel, Dietrich Klakow, and Yanai Elazar. 2023 · 2023
Later among the works it cites.
Branislav Pecher, Ivan Srba, and Maria Bielikova. 2023 · 2023
Later among the works it cites.
Is ChatGPT the ultimate data augmentation algorithm?
Frédéric Piedboeuf and Philippe Langlais. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
A ship of theseus: Curious cases of paraphrasing in llm-generated texts
Nafis Irtiza Tripto, Saranya Venkatraman, Dominik Macko, Robert Moro, Ivan Srba, Adaku Uchendu, Thai Le, and Dongwon Lee. 2023 · 2023
Later among the works it cites.
Zeroshotdataaug: Generating and augmenting training data with chatgpt
Solomon Ubani, Suleyman Olcay Polat, and Rodney Nielsen. 2023 · 2023
Later among the works it cites.
Toward learning human-aligned cross-domain robust models by countering misaligned features
Haohan Wang, Zeyi Huang, Hanlin Zhang, Yong Jae Lee, and Eric P Xing. 2022 · 2084
Closest in time.