Fetching the paper…
Reading the bibliography…
The capabilities of pretrained language models have opened opportunities to explore new application areas, but applications involving human-human interaction are limited by the fact that most data is protected from public release for privacy reasons.
An iterative design methodology for user-friendly natural language office information applications
J. F. Kelley. 1984 · 1984
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Earlier work this paper cites.
Generation of synthetic electronic medical record text
Jiaqi Guan, Runzhe Li, Sheng Yu, and Xuegong Zhang. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Data synthesis based on generative adversarial networks
Noseong Park, Mahmoud Mohammadi, Kshitij Gorde, Sushil Jajodia, Hongkyu Park, and Youngmin Kim. 2018 · 2018
Earlier work this paper cites.
Scale up event extraction learning via automatic training data generation
Ying Zeng, Yansong Feng, Rong Ma, Zheng Wang, Rui Yan, Chongde Shi, and Dongyan Zhao. 2018 · 2018
Earlier work this paper cites.
Synsys: A synthetic data generation system for healthcare applications
Jessamyn Dahmen and Diane Cook. 2019 · 2019
Earlier work this paper cites.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Feature space augmentation for long-tailed data
Peng Chu, Xiao Bian, Shaopeng Liu, and Haibin Ling. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Cited alongside, same era.
Private fl-gan: Differential privacy synthetic data generation based on federated learning
Bangzhou Xin, Wei Yang, Yangyang Geng, Sheng Chen, Shaowei Wang, and Liusheng Huang. 2020 · 2020
Cited alongside, same era.
Action-based conversations dataset: A corpus for building more in-depth task-oriented dialogue systems
Is GPT-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi. 2022 · 2022
Later among the works it cites.
LongT5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2022 · 2022
Later among the works it cites.
Generate, Annotate, and Learn: NLP with Synthetic Text
Xuanli He, Islam Nassar, Jamie Kiros, Gholamreza Haffari, and Mohammad Norouzi. 2022 · 2022
Later among the works it cites.
In-context learning for few-shot dialogue state tracking
Yushi Hu, Chia-Hsuan Lee, Tianbao Xie, Tao Yu, Noah A. Smith, and Mari Ostendorf. 2022 · 2022
Later among the works it cites.
MultiSpanQA: A dataset for multi-span question answering
Haonan Li, Martin Tomko, Maria Vasardani, and Timothy Baldwin. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Derek Chen, Howard Chen, Yi Yang, Alexander Lin, and Zhou Yu. 2021 · 2021
Cited alongside, same era.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Cited alongside, same era.
Scarecrow: A framework for scrutinizing machine text
Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Dialogue state tracking with a language model using schema-driven prompting
Chia-Hsuan Lee, Hao Cheng, and Mari Ostendorf. 2021 · 2021
Cited alongside, same era.
Annotation inconsistency and entity bias in MultiWOZ
Kun Qian, Ahmad Beirami, Zhouhan Lin, Ankita De, Alborz Geramifard, Zhou Yu, and Chinnadhurai Sankar. 2021 · 2021
Cited alongside, same era.
Synthetic data generation for grammatical error correction with tagged corruption models
Felix Stahlberg and Shankar Kumar. 2021 · 2021
Cited alongside, same era.
Human-machine collaboration approaches to build a dialogue dataset for hate speech countering
Helena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroğlu, and Marco Guerini. 2022 · 2022
Cited alongside, same era.
Soda: Million-scale dialogue distillation with social commonsense contextualization
Hyunwoo Kim, Jack Hessel, Liwei Jiang, Ximing Lu, Youngjae Yu, Pei Zhou, Ronan Le Bras, Malihe Alikhani, Gunhee Kim, Maarten Sap, et al. 2022a
Cited in the paper.
Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022a · 2022
Later among the works it cites.
Social simulacra: Creating populated prototypes for social computing systems
Joon Sung Park, Lindsay Popowski, Carrie Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2022 · 2022
Later among the works it cites.
Singan-seg: Synthetic training data generation for medical image segmentation
Vajira Thambawita, Pegah Salehi, Sajad Amouei Sheshkal, Steven A. Hicks, Hugo L. Hammer, Sravanthi Parasa, Thomas de Lange, Pål Halvorsen, and Michael A. Riegler. 2022 · 2022
Later among the works it cites.
Differentially private synthetic medical data generation using convolutional gans
Amirsina Torfi, Edward A. Fox, and Chandan K. Reddy. 2022 · 2022
Later among the works it cites.
A synthetic data generation framework for grounded dialogues
Jianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun, Yitong Li, Fei Mi, and Ruifeng Xu. 2023 · 2023
Closest in time.
Synthetic data generation with large language models for text classification: Potential and limitations
Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, and Ming Yin. 2023 · 2023
Closest in time.