Fetching the paper…
Reading the bibliography…
Data augmentation has recently seen increased interest in NLP due to more work in low-resource domains, new tasks, and the popularity of large-scale neural networks that require large amounts of training data.
A survey on face data augmentation
Xiang Wang, Kai Wang, and Shiguo Lian. 2019b · 1904
Earlier work this paper cites.
Data Augmentation for BERT Fine-Tuning in Open-Domain Question Answering
Wei Yang, Yuqing Xie, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin. 2019 · 1904
Earlier work this paper cites.
Xlda: Cross-lingual data augmentation for natural language inference and question answering
Jasdeep Singh, Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. 2019 · 1905
Earlier work this paper cites.
Data augmentation with atomic templates for spoken language understanding
Zijian Zhao, Su Zhu, and Kai Yu. 2019 · 1908
Earlier work this paper cites.
Controllable Data Synthesis Method for Grammatical Error Correction
Chencheng Wang, Liner Yang, Yun Chen, Yongping Du, and Erhong Yang. 2019a · 1909
Earlier work this paper cites.
Sequence-to-sequence Pre-training with Data Augmentation for Sentence Rewriting
Yi Zhang, Tao Ge, Furu Wei, Ming Zhou, and Xu Sun. 2019b · 1909
Earlier work this paper cites.
Transforming wikipedia into augmented data for query-focused summarization
Haichao Zhu, Li Dong, Furu Wei, Bing Qin, and Ting Liu. 2019 · 1911
Earlier work this paper cites.
Improving Robustness of Machine Translation with Synthetic Noise
Vaibhav Vaibhav, Sumeet Singh, Craig Stewart, and Graham Neubig. 2019 · 1920
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
Chris M. Bishop. 1995 · 1995
Earlier work this paper cites.
Wordnet: a lexical database for english
George A. Miller. 1995 · 1995
Earlier work this paper cites.
SMOTE: Synthetic minority over-sampling technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer. 2002 · 2002
Earlier work this paper cites.
Data Augmentation for Copy-Mechanism in Dialogue State Tracking
Xiaohui Song, Liangjun Zang, Yipeng Su, Xing Wu, Jizhong Han, and Songlin Hu. 2020 · 2002
Earlier work this paper cites.
Overview of DUC 2005
Hao T. Dang. 2005 · 2005
Earlier work this paper cites.
Data augmentation for training dialog models robust to speech recognition errors
Longshaokan Wang, Maryam Fazel-Zarandi, Aditya Tiwari, Spyros Matsoukas, and Lazaros Polymenakos. 2020 · 2006
Earlier work this paper cites.
GenERRate: Generating errors for use in grammatical error detection
Jennifer Foster and Oistein Andersen. 2009 · 2009
Earlier work this paper cites.
GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong. 2020 · 2009
Earlier work this paper cites.
How effective is task-agnostic data augmentation for pretrained transformers?
Shayne Longpre, Yu Wang, and Christopher DuBois. 2020 · 2010
Earlier work this paper cites.
Mining revision log of language learning SNS for automated Japanese error correction of second language learners
Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, and Yuji Matsumoto. 2011 · 2011
Earlier work this paper cites.
POMDP-Based Statistical Spoken Dialog Systems: A Review
S. Young, M. Gašić, B. Thomson, and J. D. Williams. 2013 · 2013
Earlier work this paper cites.
Mlsmote: Approaching imbalanced multilabel learning through synthetic instance generation
F. Charte, Antonio Rivera Rivas, María José Del Jesus, and Francisco Herrera. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification
Ellie Pavlick, Pushpendre Rastogi, Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2015 · 2015
Earlier work this paper cites.
Character-Level Convolutional Networks for Text Classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Artificial error generation for translation-based grammatical error correction
Mariano Felice. 2016 · 2016
Earlier work this paper cites.
Data recombination for neural semantic parsing
Robin Jia and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Improving Neural Machine Translation Models with Monolingual Data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Adapting neural machine translation with parallel synthetic data
Mara Chinea-Ríos, Álvaro Peris, and Francisco Casacuberta. 2017 · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor. 2017 · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
Earlier work this paper cites.
The WebNLG challenge: Generating text from RDF data
Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017 · 2017
Earlier work this paper cites.
Low-shot visual recognition by shrinking and hallucinating features
Bharath Hariharan and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Synthetic data for neural machine translation of spoken-dialects
Hany Hassan, Mostafa Elaraby, and Ahmed Tawfik. 2017 · 2017
Earlier work this paper cites.
Data augmentation for visual question answering
Kushal Kafle, Mohammed Yousefhussien, and Christopher Kanan. 2017 · 2017
Earlier work this paper cites.
Robust training under linguistic adversity
Yitong Li, Trevor Cohn, and Timothy Baldwin. 2017 · 2017
Earlier work this paper cites.
Learning to Compose Domain-Specific Transformations for Data Augmentation
AJ Ratner, HR Ehrenberg, Z Hussain, J Dunnmon, and C Ré. 2017 · 2017
Earlier work this paper cites.
Challenges in Data-to-Document Generation
Sam Wiseman, Stuart M Shieber, and Alexander M Rush. 2017 · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2017 · 2017
Earlier work this paper cites.
Using Wikipedia edits in low resource grammatical error correction
Adriane Boyd. 2018 · 2018
Earlier work this paper cites.
Autoaugment: Learning augmentation policies from data
Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le. 2018 · 2018
Earlier work this paper cites.
Counterexample-guided data augmentation
Tommaso Dreossi, Shromona Ghosh, Xiangyu Yue, Kurt Keutzer, Alberto L. Sangiovanni-Vincentelli, and Sanjit A. Seshia. 2018 · 2018
Earlier work this paper cites.
Findings of the E2E NLG challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary
Alberto Fernández, Salvador Garcia, Francisco Herrera, and Nitesh V. Chawla. 2018 · 2018
Earlier work this paper cites.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Sequence-to-sequence data augmentation for dialogue language understanding
Yutai Hou, Yijia Liu, Wanxiang Che, and Ting Liu. 2018 · 2018
Earlier work this paper cites.
Multimodal continuous emotion recognition with data augmentation using recurrent neural networks
Jian Huang, Ya Li, Jianhua Tao, Zheng Lian, Mingyue Niu, and Minghao Yang. 2018 · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
AdvEntuRe: Adversarial training for textual entailment with knowledge-guided examples
Dongyeop Kang, Tushar Khot, Ashish Sabharwal, and Eduard Hovy. 2018 · 2018
Earlier work this paper cites.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi. 2018 · 2018
Earlier work this paper cites.
Multi-source neural machine translation with data augmentation
Yuta Nishimura, Katsuhito Sudoh, Graham Neubig, and Satoshi Nakamura. 2018 · 2018
Earlier work this paper cites.
Multi-Modal Data Augmentation for End-to-end ASR
Adithya Renduchintala, Shuoyang Ding, Matthew Wiesner, and Shinji Watanabe. 2018 · 2018
Earlier work this paper cites.
Data augmentation via dependency tree morphing for low-resource languages
Gözde Gül Şahin and Mark Steedman. 2018 · 2018
Earlier work this paper cites.
δ \delta -encoder: an effective sample synthesis method for few-shot object recognition
Eli Schwartz, Leonid Karlinsky, Joseph Shtok, Sivan Harary, Mattias Marder, Abhishek Kumar, Rogerio Feris, Raja Giryes, and Alex M Bronstein. 2018 · 2018
Earlier work this paper cites.
TNT-NLG, System 2: Data repetition and meaning representation manipulation to improve neural generation
Shubhangi Tandon, TS Sharath, Shereen Oraby, Lena Reed, Stephanie Lukin, and Marilyn Walker. 2018 · 2018
Earlier work this paper cites.
SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018a · 2018
Earlier work this paper cites.
SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018b · 2018
Cited alongside, same era.
Low resource multi-modal data augmentation for end-to-end ASR
Matthew Wiesner, Adithya Renduchintala, Shinji Watanabe, Chunxi Liu, Najim Dehak, and Sanjeev Khudanpur. 2018 · 2018
Cited alongside, same era.
Noising and Denoising Natural Language: Diverse Backtranslation for Grammar Correction
Ziang Xie, Guillaume Genthial, Stanley Xie, Andrew Ng, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
Augmenting Image Question Answering Dataset by Exploiting Image Captions
Masashi Yokota and Hideki Nakayama. 2018 · 2018
Cited alongside, same era.
Fast and accurate reading comprehension by combining self-attention and convolution
Adams Wei Yu, David Dohan, Quoc Le, Thang Luong, Rui Zhao, and Kai Chen. 2018 · 2018
Cited alongside, same era.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2020 · 2020
Later among the works it cites.
An analysis of simple data augmentation for named entity recognition
Xiang Dai and Heike Adel. 2020 · 2020
Later among the works it cites.
DAGA: Data augmentation with a generation approach forLow-resource tagging tasks
Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, and Chunyan Miao. 2020 · 2020
Later among the works it cites.
Syntax-aware data augmentation for neural machine translation
Sufeng Duan, Hai Zhao, Dongdong Zhang, and Rui Wang. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
On adversarial mixup resynthesis
Christopher Beckham, Sina Honari, Vikas Verma, Alex M. Lamb, Farnoosh Ghadiri, R Devon Hjelm, Yoshua Bengio, and Chris Pal. 2019 · 2019
Cited alongside, same era.
Neural fuzzy repair: Integrating fuzzy matches into neural machine translation
Bram Bulte and Arda Tezcan. 2019 · 2019
Cited alongside, same era.
A neural grammatical error correction system built on better pre-training and sequential transfer learning
Yo Joong Choe, Jiyeon Ham, Kyubyong Park, and Yeoil Yoon. 2019 · 2019
Cited alongside, same era.
A kernel theory of modern data augmentation
Tri Dao, Albert Gu, Alexander J. Ratner, Virginia Smith, Christopher De Sa, and Christopher Ré. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Keep calm and switch on! Preserving sentiment and fluency in semantic text exchange
Steven Y. Feng, Aaron W. Li, and Jesse Hoey. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Alexander R. Fabbri, Simeng Han, Haoyuan Li, Haoran Li, Marjan Ghazvininejad, Shafiq Joty, Dragomir Radev, and Yashar Mehdad. 2020 · 2020
Later among the works it cites.
Data augmentation techniques for the video question answering task
Alex Falcon, Oswald Lanz, and Giuseppe Serra. 2020 · 2020
Later among the works it cites.
Patchup: A regularization technique for convolutional neural networks
Mojtaba Faramarzi, Mohammad Amini, Akilesh Badrinaaraayanan, Vikas Verma, and Sarath Chandar. 2020 · 2020
Later among the works it cites.
GenAug: Data augmentation for finetuning text generators
Steven Y. Feng, Varun Gangal, Dongyeop Kang, Teruko Mitamura, and Eduard Hovy. 2020 · 2020
Later among the works it cites.
Paraphrase augmented task-oriented dialog generation
Silin Gao, Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020 · 2020
Later among the works it cites.
Simple copy-paste is a strong data augmentation method for instance segmentation
Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D. Cubuk, Quoc V. Le, and Barret Zoph. 2020 · 2020
Later among the works it cites.
Tradeoffs in data augmentation: An empirical study
Raphael Gontijo-Lopes, Sylvia J. Smullin, Ekin D. Cubuk, and Ethan Dyer. 2020 · 2020
Later among the works it cites.
Sequence-level mixed sample data augmentation
Demi Guo, Yoon Kim, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Nonlinear mixup: Out-of-manifold data augmentation for text classification
Hongyu Guo. 2020 · 2020
Later among the works it cites.
Data augmentation instead of explicit regularization
Alex Hernández-García and Peter König. 2020 · 2020
Later among the works it cites.
TaPas: Weakly supervised table parsing via pre-training
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020 · 2020
Later among the works it cites.
An empirical survey of data augmentation for time series classification with neural networks
Brian Kenji Iwana and Seiichi Uchida. 2020 · 2020
Later among the works it cites.
Quantifying the evaluation of heuristic methods for textual data augmentation
Omid Kashefi and Rebecca Hwa. 2020 · 2020
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Later among the works it cites.
A syntactic rule-based framework for parallel data synthesis in japanese gec
Alex Kimn. 2020 · 2020
Later among the works it cites.
Syntax-guided controlled generation of paraphrases
Ashutosh Kumar, Kabir Ahuja, Raghuram Vadapalli, and Partha Talukdar. 2020 · 2020
Later among the works it cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
A survey of text data augmentation
Pei Liu, Xuemin Wang, Chao Xiang, and Weiye Meng. 2020a · 2020
Later among the works it cites.
Data boost: Text data augmentation through reinforcement learning guided conditional generation
Ruibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma, Lili Wang, and Soroush Vosoughi. 2020b · 2020
Later among the works it cites.
Simple is better! lightweight data augmentation for low resource slot filling and intent classification
Samuel Louvan and Bernardo Magnini. 2020 · 2020
Later among the works it cites.
Gender Bias in Neural Natural Language Processing , pages 189–202. Springer International Publishing, Cham
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020 · 2020
Later among the works it cites.
Denoising pre-training and data augmentation strategies for enhanced RDF verbalization with transformers
Sebastien Montella, Betty Fabre, Tanguy Urvoy, Johannes Heinecke, and Lina Rojas-Barahona. 2020 · 2020
Later among the works it cites.
Improving robustness by augmenting training sentences with predicate-argument structures
Nafise Sadat Moosavi, Marcel de Boer, Prasetya Ajie Utama, and Iryna Gurevych. 2020 · 2020
Later among the works it cites.
Multimodal dialogue state tracking by qa approach with data augmentation
Xiangyang Mou, Brandyn Sigouin, Ian Steenstra, and Hui Su. 2020 · 2020
Later among the works it cites.
SSMBA: Self-supervised manifold based data augmentation for improving out-of-domain robustness
Nathan Ng, Kyunghyun Cho, and Marzyeh Ghassemi. 2020 · 2020
Later among the works it cites.
Data diversification: A simple strategy for neural machine translation
Xuan-Phi Nguyen, Shafiq Joty, Kui Wu, and Ai Ti Aw. 2020 · 2020
Later among the works it cites.
Named entity recognition for social media texts with semantic augmentation
Yuyang Nie, Yuanhe Tian, Xiang Wan, Yan Song, and Bo Dai. 2020 · 2020
Later among the works it cites.
Dictionary-based data augmentation for cross-domain neural machine translation
Wei Peng, Chongxuan Huang, Tianhao Li, Yun Chen, and Qun Liu. 2020 · 2020
Later among the works it cites.
Cosda-ml: Multi-lingual code-switching data augmentation for zero-shot cross-lingual nlp
Libo Qin, Minheng Ni, Yue Zhang, and Wanxiang Che. 2020 · 2020
Later among the works it cites.
Textual data augmentation for efficient active learning on tiny datasets
Husam Quteineh, Spyridon Samothrakis, and Richard Sutcliffe. 2020 · 2020
Later among the works it cites.
Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering
Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, and Jacopo Staiano. 2020 · 2020
Later among the works it cites.
Semantic equivalent adversarial data augmentation for visual question answering
Ruixue Tang, Chao Ma, Wei Emma Zhang, Qi Wu, and Xiaokang Yang. 2020 · 2020
Later among the works it cites.
Improving Grammatical Error Correction with Data Augmentation by Editing Latent Representation
Zhaohong Wan, Xiaojun Wan, and Wenguang Wang. 2020 · 2020
Later among the works it cites.
A Comparative Study of Synthetic Data Generation Methods for Grammatical Error Correction
Max White and Alla Rozovskaya. 2020 · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Later among the works it cites.
MDA: Multimodal Data Augmentation Framework for Boosting Performance on Image-Text Sentiment/Emotion Classification Tasks
N. Xu, W. Mao, P. Wei, and D. Zeng. 2020 · 2020
Later among the works it cites.
G-daug: Generative data augmentation for commonsense reasoning
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020 · 2020
Later among the works it cites.
Dialog State Tracking with Reinforced Data Augmentation
Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020 · 2020
Later among the works it cites.
SeqMix: Augmenting Active Sequence Labeling via Sequence Mixup
Rongzhi Zhang, Yue Yu, and Chao Zhang. 2020 · 2020
Later among the works it cites.
Grounded adaptation for zero-shot executable semantic parsing
Victor Zhong, Mike Lewis, Sida I. Wang, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Nareor: The narrative reordering problem
Varun Gangal, Steven Y. Feng, Eduard Hovy, and Teruko Mitamura. 2021 · 2021
Closest in time.
Conversation graph: Data augmentation, training, and evaluation for non-deterministic dialogue management
Milan Gritta, Gerasimos Lampouras, and Ignacio Iacobacci. 2021 · 2021
Closest in time.
Explaining the efficacy of counterfactually augmented data
Divyansh Kaushik, Amrith Setlur, Eduard H. Hovy, and Zachary Chase Lipton. 2021 · 2021
Closest in time.
Neural data augmentation via example extrapolation
Kenton Lee, Kelvin Guu, Luheng He, Tim Dozat, and Hyung Won Chung. 2021 · 2021
Closest in time.
Data augmentation for abstractive query-focused multi-document summarization
Ramakanth Pasunuru, Asli Celikyilmaz, Michel Galley, Chenyan Xiong, Yizhe Zhang, Mohit Bansal, and Jianfeng Gao. 2021 · 2021
Closest in time.
Substructure Substitution: Structured Data Augmentation for NLP
Haoyue Shi, Karen Livescu, and Kevin Gimpel. 2021 · 2021
Closest in time.
Negative data augmentation
Abhishek Sinha, Kumar Ayush, Jiaming Song, Burak Uzkent, Hongxia Jin, and Stefano Ermon. 2021 · 2021
Closest in time.
Nandan Thakur, Nils Reimers, Johannes Daxenberger, and Iryna Gurevych. 2021 · 2021
Closest in time.
Revisiting Recurrent Networks for Paraphrastic Sentence Embeddings
John Wieting and Kevin Gimpel. 2017 · 2088
Closest in time.