Fetching the paper…
Reading the bibliography…
We study a family of data augmentation methods, substructure substitution (SUB2), for natural language processing (NLP) tasks.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. 2019 · 1904
Earlier work this paper cites.
Augmenting data with mixup for sentence classification: An empirical study
Hongyu Guo, Yongyi Mao, and Richong Zhang. 2019 · 1905
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Mathematical and computational aspects of lexicalized grammars
Yves Schabes. 1990 · 1990
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Semi-supervised semantic role labeling
Hagen Fürstenau and Mirella Lapata. 2009 · 2009
Earlier work this paper cites.
The nxt-format switchboard corpus: a rich resource for investigating the syntax, semantics, pragmatics and prosody of dialogue
Sasha Calhoun, Jean Carletta, Jason M Brenier, Neil Mayo, Dan Jurafsky, Mark Steedman, and David Beaver. 2010 · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Foreebank: Syntactic analysis of customer support forums
Rasoul Kaljahi, Jennifer Foster, Johann Roturier, Corentin Ribeyre, Teresa Lynn, and Joseph Le Roux. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
That’s so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using #petpeeve tweets
William Yang Wang and Diyi Yang. 2015 · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Long short-term memory over recursive structures
Xiaodan Zhu, Parinaz Sobihani, and Hongyu Guo. 2015 · 2015
Earlier work this paper cites.
Data recombination for neural semantic parsing
Robin Jia and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Universal Dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Improved relation classification by deep recurrent neural networks with data augmentation
Yan Xu, Ran Jia, Lili Mou, Ge Li, Yunchuan Chen, Yangyang Lu, and Zhi Jin. 2016 · 2016
Earlier work this paper cites.
Training data augmentation for low-resource morphological inflection
Toms Bergmanis, Katharina Kann, Hinrich Schütze, and Sharon Goldwater. 2017 · 2017
Earlier work this paper cites.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
Earlier work this paper cites.
Data augmentation for visual question answering
Kushal Kafle, Mohammed Yousefhussien, and Christopher Kanan. 2017 · 2017
Cited alongside, same era.
Data augmentation for morphological reinflection
Miikka Silfverberg, Adam Wiemerslage, Ling Liu, and Lingshuang Jack Mao. 2017 · 2017
Cited alongside, same era.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Cited alongside, same era.
Improved sentence modeling using suffix bidirectional lstm
Siddhartha Brahma. 2018 · 2018
Cited alongside, same era.
Sequence-to-sequence data augmentation for dialogue language understanding
Generating fluent adversarial examples for natural languages
Huangzhao Zhang, Hao Zhou, Ning Miao, and Lei Li. 2019 · 2019
Later among the works it cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Later among the works it cites.
Good-enough compositional data augmentation
Jacob Andreas. 2020 · 2020
Later among the works it cites.
Logic-guided data augmentation and regularization for consistent question answering
Akari Asai and Hannaneh Hajishirzi. 2020 · 2020
Later among the works it cites.
Local additivity based data augmentation for semi-supervised NER
Jiaao Chen, Zhenghui Wang, Ran Tian, Zichao Yang, and Diyi Yang. 2020 · 2020
Later among the works it cites.
An analysis of simple data augmentation for named entity recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yutai Hou, Yijia Liu, Wanxiang Che, and Ting Liu. 2018 · 2018
Cited alongside, same era.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein. 2018 · 2018
Cited alongside, same era.
Contextual augmentation: Data augmentation by words with paradigmatic relations
Sosuke Kobayashi. 2018 · 2018
Cited alongside, same era.
Style transfer through back-translation
Shrimai Prabhumoye, Yulia Tsvetkov, Ruslan Salakhutdinov, and Alan W Black. 2018 · 2018
Cited alongside, same era.
Data augmentation via dependency tree morphing for low-resource languages
Gözde Gül Şahin and Mark Steedman. 2018 · 2018
Cited alongside, same era.
On tree-based neural sentence modeling
Haoyue Shi, Hao Zhou, Jiaze Chen, and Lei Li. 2018b · 2018
Cited alongside, same era.
SwitchOut: an efficient data augmentation algorithm for neural machine translation
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Xiang Dai and Heike Adel. 2020 · 2020
Later among the works it cites.
Data augmentation via subtree swapping for dependency parsing of low-resource languages
Mathieu Dehouck and Carlos Gómez-Rodríguez. 2020 · 2020
Later among the works it cites.
DAGA: Data augmentation with a generation approach for low-resource tagging tasks
Bosheng Ding, Linlin Liu, Lidong Bing, Canasai Kruengkrai, Thien Hai Nguyen, Shafiq Joty, Luo Si, and Chunyan Miao. 2020 · 2020
Later among the works it cites.
Sequence-level mixed sample data augmentation
Demi Guo, Yoon Kim, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Controllable meaning representation to text generation: Linearization and data augmentation strategies
Chris Kedzie and Kathleen McKeown. 2020 · 2020
Later among the works it cites.
Tell me how to ask again: Question data augmentation with controllable rewriting in continuous space
Dayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan, Jiusheng Chen, Jiancheng Lv, Nan Duan, and Ming Zhou. 2020a · 2020
Later among the works it cites.
Data boost: Text data augmentation through reinforcement learning guided conditional generation
Ruibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma, Lili Wang, and Soroush Vosoughi. 2020b · 2020
Later among the works it cites.
Gender bias in neural natural language processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2020 · 2020
Later among the works it cites.
Syntactic data augmentation increases robustness to inference heuristics
Junghyun Min, R. Thomas McCoy, Dipanjan Das, Emily Pitler, and Tal Linzen. 2020 · 2020
Later among the works it cites.
Rethinking self-attention: Towards interpretability in neural parsing
Khalil Mrini, Franck Dernoncourt, Quan Hung Tran, Trung Bui, Walter Chang, and Ndapa Nakashole. 2020 · 2020
Later among the works it cites.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Later among the works it cites.
Textual data augmentation for efficient active learning on tiny datasets
Husam Quteineh, Spyridon Samothrakis, and Richard Sutcliffe. 2020 · 2020
Later among the works it cites.
On the role of supervision in unsupervised constituency parsing
Haoyue Shi, Karen Livescu, and Kevin Gimpel. 2020 · 2020
Later among the works it cites.
Mixup-transformer: Dynamic data augmentation for NLP tasks
Lichao Sun, Congying Xia, Wenpeng Yin, Tingting Liang, Philip Yu, and Lifang He. 2020 · 2020
Later among the works it cites.
Improving grammatical error correction with data augmentation by editing latent representation
Zhaohong Wan, Xiaojun Wan, and Wenguang Wang. 2020 · 2020
Later among the works it cites.
Generative data augmentation for commonsense reasoning
Yiben Yang, Chaitanya Malaviya, Jared Fernandez, Swabha Swayamdipta, Ronan Le Bras, Ji-Ping Wang, Chandra Bhagavatula, Yejin Choi, and Doug Downey. 2020 · 2020
Later among the works it cites.
Variational hierarchical dialog autoencoder for dialog state tracking data augmentation
Kang Min Yoo, Hanbit Lee, Franck Dernoncourt, Trung Bui, Walter Chang, and Sang-goo Lee. 2020 · 2020
Later among the works it cites.