2023

ZeroShotDataAug: Generating and Augmenting Training Data with ChatGPT

Ubani, Solomon, Polat, Suleyman Olcay, Nielsen, Rodney

Understand

In this paper, we investigate the use of data obtained from prompting a large generative language model, ChatGPT, to generate synthetic training data with the aim of augmenting data in low resource scenarios.

  • We show that with appropriate task-specific ChatGPT prompts, we outperform the most popular existing approaches for such data augmentation.
  • Furthermore, we investigate methodologies for evaluating the similarity of the augmented data generated from ChatGPT with the aim of validating and assessing the quality of the data generated.

Built on

Similar

Then

  • “Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers”

    Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang and Ming Zhou · 2020

    Later among the works it cites.

  • “Chataug: Leveraging chatgpt for text data augmentation”

    Original

    Haixing Dai, Zhengliang Liu, Wenxiong Liao, Xiaoke Huang, Zihao Wu, Lin Zhao, Wei Liu, Ninghao Liu, Sheng Li and Dajiang Zhu · 2023

    Closest in time.

  • “GPT-3.5” [Accessed: 07.04.2023]

    OpenAI · 2023

    Closest in time.

  • “Introducing ChatGPT.” [Accessed: 07.04.2023]

    OpenAI · 2023

    Closest in time.

  • “SentenceTransformers Documentation” [Accessed: 07.04.2023]

    SBERT.net · 2023

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…