Fetching the paper…
Reading the bibliography…
Data augmentation is an important component in the robustness evaluation of models in natural language processing (NLP) and in enhancing the diversity of the data they are trained on.
Generating textual adversarial examples for deep learning models: A survey
Wei Emma Zhang, Quan Z. Sheng, and Ahoud Abdulrahmn F. Alhazmi. 2019a · 1901
Earlier work this paper cites.
Text processing like humans do: Visually attacking and shielding NLP systems
Steffen Eger, Gözde Gül Sahin, Andreas Rücklé, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. 2019b · 1903
Earlier work this paper cites.
Eliminet: A model for eliminating options for reading comprehension with multiple choice questions
Soham Parikh, Ananya B. Sai, Preksha Nema, and Mitesh M. Khapra. 2019 · 1904
Earlier work this paper cites.
Simple bert models for relation extraction and semantic role labeling
Peng Shi and Jimmy Lin. 2019a · 1904
Earlier work this paper cites.
Simple BERT models for relation extraction and semantic role labeling
Peng Shi and Jimmy Lin. 2019b · 1904
Earlier work this paper cites.
Naver labs europe’s systems for the wmt19 machine translation robustness task
Alexandre Berard, Ioan Calapodescu, and Claude Roux. 2019 · 1907
Earlier work this paper cites.
Vox populi (the wisdom of crowds)
Francis Galton. 1907 · 1907
Earlier work this paper cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton. 2019 · 1909
Earlier work this paper cites.
Commongen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2019 · 1911
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Designing Statistical Language Learners: Experiments on Noun Compounds
Mark Lauer. 1995 · 1995
Earlier work this paper cites.
WordNet: An electronic lexical database
George A Miller. 1998 · 1998
Earlier work this paper cites.
Cluener2020: Fine-grained name entity recognition for chinese
Liang Xu, Qianqian Dong, Cong Yu, Yin Tian, Weitang Liu, Lu Li, and Xuanwei Zhang. 2020 · 2001
Earlier work this paper cites.
The necessity of parsing for predicate argument recognition
Daniel Gildea and Martha Stone Palmer. 2002 · 2002
Earlier work this paper cites.
From treebank to propbank
Paul R. Kingsbury and Martha Palmer. 2002 · 2002
Earlier work this paper cites.
Kyubyong Park and Seanie Lee. 2020 · 2004
Earlier work this paper cites.
Correcting spelling errors by modelling their causes
Sebastian Deorowicz and Marcin G Ciura. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The proposition bank: An annotated corpus of semantic roles
Martha Palmer, Paul R. Kingsbury, and Daniel Gildea. 2005 · 2005
Earlier work this paper cites.
Respectful Disability Language: Here’s What’s Up!
2006 · 2006
Earlier work this paper cites.
Nltk: the natural language toolkit
Steven Bird. 2006 · 2006
Earlier work this paper cites.
Text data augmentation: Towards better detection of spear-phishing emails
Mehdi Regina, Maxime Meyer, and Sébastien Goutal. 2020 · 2007
Earlier work this paper cites.
An overview of the tesseract OCR engine
R. Smith. 2007 · 2007
Earlier work this paper cites.
Assessing demographic bias in named entity recognition
Shubhanshu Mishra, Sijun He, and Luca Belli. 2020 · 2008
Earlier work this paper cites.
Easily identifiable discourse relations
Emily Pitler, Mridhula Raghupathy, Hena Mehta, Ani Nenkova, Alan Lee, and Aravind Joshi. 2008 · 2008
Earlier work this paper cites.
The Penn Discourse TreeBank 2.0
Rashmi Prasad, Nikhil Dinesh, Alan Lee, Eleni Miltsakaki, Livio Robaldo, Aravind Joshi, and Bonnie Webber. 2008 · 2008
Earlier work this paper cites.
Contextualized perturbation for textual adversarial attack
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2020a · 2009
Earlier work this paper cites.
Contextualized perturbation for textual adversarial attack
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2020b · 2009
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Çelebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. 2020 · 2010
Earlier work this paper cites.
Wisdom of the crowds in minimum spanning tree problems
Sheng Kung Yi, Mark Steyvers, Michael Lee, and Matthew Dry. 2010 · 2010
Earlier work this paper cites.
Improved semantic role labeling using parameterized neighborhood memory adaptation
Ishan Jindal, Ranit Aharonov, Siddhartha Brahma, Huaiyu Zhu, and Yunyao Li. 2020 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
English propbank annotation guidelines
Claire Bonial, Jena Hwang, Julia Bonn, Kathryn Conger, Olga Babko-Malaya, and Martha Palmer. 2012 · 2012
Earlier work this paper cites.
What is a paraphrase?
Rahul Bhagat and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
Semeval-2013 task 4: Free paraphrases of noun compounds
Iris Hendrickx, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Stan Szpakowicz, and Tony Veale. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Um… who like says you know: Filler word use as a function of age, gender, and personality
Charlyn M Laserna, Yi-Tai Seih, and James W Pennebaker. 2014 · 2014
Earlier work this paper cites.
emoji2vec: Learning emoji representations from their description
Ben Eisner, Tim Rocktäschel, Isabelle Augenstein, Matko Bošnjak, and Sebastian Riedel. 2016 · 2016
Earlier work this paper cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
Ashwin K. Vijayakumar, Michael Cogswell, Ramprasaath R. Selvaraju, Qing Sun, Stefan Lee, David J. Crandall, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
Natural language processing for the long tail
David Bamman. 2017 · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
Earlier work this paper cites.
Neural network methods for natural language processing
Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
Pushing the limits of paraphrastic sentence embeddings with millions of machine translations
John Wieting and Kevin Gimpel. 2017 · 2017
Cited alongside, same era.
Learning paraphrastic sentence embeddings from back-translated bitext
John Wieting, Jonathan Mallinson, and Kevin Gimpel. 2017 · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017 · 2017
Cited alongside, same era.
The narrativeqa reading comprehension challenge
Tomáš Kočiskỳ, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Cited alongside, same era.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Cited alongside, same era.
Participatory research for low-resourced machine translation: A case study in African languages
Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020 · 2020
Later among the works it cites.
Looking inside noun compounds: Unsupervised prepositional and free paraphrasing using language models
Girishkumar Ponkiya, Rudra Murthy, Pushpak Bhattacharyya, and Girish Palshikar. 2020 · 2020
Later among the works it cites.
Cosda-ml: Multi-lingual code-switching data augmentation for zero-shot cross-lingual nlp
Libo Qin, Minheng Ni, Yue Zhang, and Wanxiang Che. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Higher-order coreference resolution with coarse-to-fine inference
Kenton Lee, Luheng He, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Content preserving text generation with attribute controls
Lajanugen Logeswaran, Honglak Lee, and Samy Bengio. 2018 · 2018
Cited alongside, same era.
Treat us like the sequences we are: Prepositional paraphrasing of noun compounds using lstm
Girishkumar Ponkiya, Kevin Patel, Pushpak Bhattacharyya, and Girish Palshikar. 2018 · 2018
Cited alongside, same era.
Paraphrase to explicate: Revealing implicit noun-compound relations
Vered Shwartz and Ido Dagan. 2018 · 2018
Cited alongside, same era.
Diverse beam search for improved description of complex scenes
Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018 · 2018
Cited alongside, same era.
Text processing like humans do: Visually attacking and shielding NLP systems
Steffen Eger, Gözde Gül Şahin, Andreas Rücklé, Ji-Ung Lee, Claudia Schulz, Mohsen Mesgar, Krishnkant Swarnkar, Edwin Simpson, and Iryna Gurevych. 2019a · 2019
Cited alongside, same era.
Studying cultural differences in emoji usage across the east and the west
Sharath Chandra Guntuku, Mingyang Li, Louis Tay, and Lyle H Ungar. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
It’s morphin’ time! Combating linguistic discrimination with inflectional perturbations
Samson Tan, Shafiq Joty, Min-Yen Kan, and Richard Socher. 2020 · 2020
Later among the works it cites.
OPUS-MT — Building open translation services for the World
Jörg Tiedemann and Santhosh Thottingal. 2020 · 2020
Later among the works it cites.
Urban dictionary embeddings for slang NLP applications
Steven Wilson, Walid Magdy, Barbara McGillivray, Kiran Garimella, and Gareth Tyson. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Later among the works it cites.
Frequently misspelled word list for dyslexia
Smorga’s Board. 2021 · 2021
Closest in time.
An empirical survey of data augmentation for limited data learning in nlp
Jiaao Chen, Derek Tam, Colin Raffel, Mohit Bansal, and Diyi Yang. 2021 · 2021
Closest in time.
Protaugment: Unsupervised diverse short-texts paraphrasing for intent detection meta-learning
Thomas Dopierre, Christophe Gravier, and Wilfried Logerais. 2021 · 2021
Closest in time.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, et al. 2021 · 2021
Closest in time.
A survey of data augmentation approaches for nlp
Steven Y Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021 · 2021
Closest in time.
Nareor: The narrative reordering problem
Varun Gangal, Steven Y Feng, Eduard Hovy, and Teruko Mitamura. 2021 · 2021
Closest in time.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Closest in time.
Robustness Gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Rajani, Jesse Vig, Samson Tan, Jason Wu, Stephan Zheng, Caiming Xiong annd Mohit Bansal, and Christopher Ré. 2021 · 2021
Closest in time.
Candle: Decomposing conditional and conjunctive queries for task-oriented dialogue systems
Aadesh Gupta, Kaustubh D. Dhole, Rahul Tarway, Swetha Prabhakar, and Ashish Shrivastava. 2021 · 2021
Closest in time.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Closest in time.
Can vectors read minds better than experts? comparing data augmentation strategies for the automated scoring of children’s mindreading ability
Venelin Kovatchev, Phillip Smith, Mark Lee, and Rory Devine. 2021 · 2021
Closest in time.
Continual mixed-language pre-training for extremely low-resource neural machine translation
Zihan Liu, Genta Indra Winata, and Pascale Fung. 2021 · 2021
Closest in time.
Automatic construction of evaluation suites for natural language generation datasets
Simon Mille, Kaustubh D. Dhole, Saad Mahamood, Laura Perez-Beltrachini, Varun Gangal, Mihir Kale, Emiel van Miltenburg, and Sebastian Gehrmann. 2021 · 2021
Closest in time.
Empirical error modeling improves robustness of noisy neural sequence labeling
Marcin Namysl, Sven Behnke, and Joachim Köhler. 2021 · 2021
Closest in time.
Data augmentation by concatenation for low-resource translation: A mystery and a solution
Toan Q Nguyen, Kenton Murray, and David Chiang. 2021 · 2021
Closest in time.
Transformers Interpret
Charles Pierse. 2021 · 2021
Closest in time.
Does robustness improve fairness? approaching fairness with word substitution robustness methods for text classification
Yada Pruksachatkun, Satyapriya Krishna, Jwala Dhamala, Rahul Gupta, and Kai-Wei Chang. 2021 · 2021
Closest in time.
WGND 2.0
Julio Raffo. 2021 · 2021
Closest in time.
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021 · 2021
Closest in time.
NoiseQA: Challenge Set Evaluation for User-Centric Question Answering
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, and Alan W Black. 2021 · 2021
Closest in time.
Substructure substitution: Structured data augmentation for NLP
Haoyue Shi, Karen Livescu, and Kevin Gimpel. 2021 · 2021
Closest in time.
Saying No is An Art: Contextualized Fallback Responses for Unanswerable Dialogue Queries
Ashish Shrivastava, Kaustubh Dhole, Abhinav Bhatt, and Sharvani Raghunath. 2021 · 2021
Closest in time.
Better robustness by more coverage: Adversarial and mixup data augmentation for robust finetuning
Chenglei Si, Zhengyan Zhang, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Qun Liu, and Maosong Sun. 2021 · 2021
Closest in time.
They, them, theirs: Rewriting with gender-neutral english
Tony Sun, Kellie Webster, Apurva Shah, William Yang Wang, and Melvin Johnson. 2021 · 2021
Closest in time.
Code-mixing on sesame street: Dawn of the adversarial polyglots
Samson Tan and Shafiq Joty. 2021 · 2021
Closest in time.
A closer look into the robustness of neural dependency parsers using better adversarial examples
Yuxuan Wang, Wanxiang Che, Ivan Titov, Shay B. Cohen, Zhilin Lei, and Ting Liu. 2021b · 2021
Closest in time.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2021 · 2021
Closest in time.
Data augmentation for low-resource named entity recognition using backtranslation
Usama Yaseen and Stefan Langer. 2021 · 2021
Closest in time.
Smat: An attention-based deep learning solution to the automation of schema matching
Jing Zhang, Bonggun Shin, Jinho D Choi, and Joyce C Ho. 2021 · 2021
Closest in time.
Gemv2: Multilingual nlg benchmarking in a single line of code
Sebastian Gehrmann, Abhik Bhattacharjee, Abinaya Mahendiran, Alex Wang, Alexandros Papangelis, Aman Madaan, Angelina McMillan-Major, Anna Shvets, Ashish Upadhyay, Bingsheng Yao, Bryan Wilie, Chandra Bhagavatula, Chaobin You, Craig Thomson, Cristina Garbacea, Dakuo Wang, Daniel Deutsch, Deyi Xiong, Di Jin, Dimitra Gkatzia, Dragomir Radev, Elizabeth Clark, Esin Durmus, Faisal Ladhak, Filip Ginter, Genta Indra Winata, Hendrik Strobelt, Hiroaki Hayashi, Jekaterina Novikova, Jenna Kanerva, Jenny Chim, Jiawei Zhou, Jordan Clive, Joshua Maynez, João Sedoc, Juraj Juraska, Kaustubh Dhole, Khyathi Raghavi Chandu, Laura Perez-Beltrachini, Leonardo F. R. Ribeiro, Lewis Tunstall, Li Zhang, Mahima Pushkarna, Mathias Creutz, Michael White, Mihir Sanjay Kale, Moussa Kamal Eddine, Nico Daheim, Nishant Subramani, Ondrej Dusek, Paul Pu Liang, Pawan Sasanka Ammanamanchi, Qi Zhu, Ratish Puduppully, Reno Kriz, Rifat Shahriyar, Ronald Cardenas, Saad Mahamood, Salomey Osei, Samuel Cahyawijaya, Sanja Štajner, Sebastien Montella, Shailza, Shailza Jolly, Simon Mille, Tahmid Hasan, Tianhao Shen, Tosin Adewumi, Vikas Raunak, Vipul Raheja, Vitaly Nikolaev, Vivian Tsai, Yacine Jernite, Ying Xu, Yisi Sang, Yixin Liu, and Yufang Hou. 2022 · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022 · 2022
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017a · 2031
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017b · 2031
Closest in time.