Fetching the paper…
Reading the bibliography…
This paper focuses on the data augmentation for low-resource NLP tasks where the training set is limited.
Eda: Easy data augmentation techniques for boosting performance on text classification tasks
Jason Wei and Kai Zou · 1901
Earlier work this paper cites.
Reasoning over paragraph effects in situations
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner · 1908
Earlier work this paper cites.
Quartz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark · 1909
Earlier work this paper cites.
Commongen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren · 1911
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Sang and Fien De Meulder · 2003
Earlier work this paper cites.
SemEval-2019 task 3: EmoContext contextual emotion detection in text
Ankush Chatterjee, Kedhar Nath Narahari, Meghana Joshi, and Puneet Agrawal · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
Alec Go, Richa Bhayani, and Lei Huang · 2009
Earlier work this paper cites.
Contributions to the study of sms spam filtering: new collection and results
Tiago A Almeida, José María G Hidalgo, and Akebo Yamakami · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports
Harsha Gurulingappa, Abdul Mateen Rajput, Angus Roberts, Juliane Fluck, Martin Hofmann-Apitius, and Luca Toldo · 2012
Earlier work this paper cites.
English-asl gloss parallel corpus 2012: Aslg-pc12
Achraf Othman and Mohamed Jemni · 2012
Earlier work this paper cites.
Resolving complex cases of definite pronouns: the winograd schema challenge
Altaf Rahman and Vincent Ng · 2012
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala · 2014
Earlier work this paper cites.
A sick cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Question-answer driven semantic role labeling: Using natural language to annotate natural language
Luheng He, Mike Lewis, and Luke Zettlemoyer · 2015
Earlier work this paper cites.
Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, Sören Auer, et al · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston · 2015
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Earlier work this paper cites.
That’s so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using# petpeeve tweets
William Yang Wang and Diyi Yang · 2015
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Generating text from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli · 2016
Earlier work this paper cites.
Semeval-2016 task 6: Detecting stance in tweets
Saif Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Deepstance at semeval-2016 task 6: Detecting stance in tweets using character and word-level cnns
Prashanth Vijayaraghavan, Ivan Sysoev, Soroush Vosoughi, and Deb Roy · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber · 2017
Earlier work this paper cites.
Searchqa: A new q&a dataset augmented with context from a search engine
Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz · 2017
Earlier work this paper cites.
Android apps and user feedback: a dataset for software evolution and quality improvement
Giovanni Grano, Andrea Di Sorbo, Francesco Mercaldo, Corrado A Visaggio, Gerardo Canfora, and Sebastiano Panichella · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
” liar, liar pants on fire”: A new benchmark dataset for fake news detection
William Yang Wang · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Johannes Welbl, Nelson F Liu, and Matt Gardner · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Victor Zhong, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Learning to split and rephrase from wikipedia edit history
Jan A Botha, Manaal Faruqui, John Alex, Jason Baldridge, and Dipanjan Das · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Cited alongside, same era.
Hate speech dataset from a white supremacy forum
Ona De Gibert, Naiara Perez, Aitor García-Pablos, and Montse Cuadros · 2018
Cited alongside, same era.
Identifying Well-formed Natural Language Questions
Manaal Faruqui and Dipanjan Das · 2018
Cited alongside, same era.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark · 2018
Cited alongside, same era.
Abstractive summarization of reddit posts with multi-level memory networks
Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim · 2018
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth · 2019
Later among the works it cites.
Do not have enough data? deep learning to the rescue!
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, Naama Tepper, and Naama Zwerdling · 2020
Later among the works it cites.
Beat the AI: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp · 2020
Later among the works it cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al · 2020
Later among the works it cites.
Mocha: A dataset for training and evaluating generative reading comprehension metrics
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Cited alongside, same era.
DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension
Amrita Saha, Rahul Aralikatte, Mitesh M. Khapra, and Karthik Sankaranarayanan · 2018
Cited alongside, same era.
Carer: Contextualized affect representations for emotion recognition
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen · 2018
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Cited alongside, same era.
Onestopenglish corpus: A new corpus for automatic readability assessment and text simplification
Sowmya Vajjala and Ivana Lučić · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Cited alongside, same era.
Later among the works it cites.
An analysis of simple data augmentation for named entity recognition
Xiang Dai and Heike Adel · 2020
Later among the works it cites.
Climate-fever: A dataset for verification of real-world climate claims, 2020
Thomas Diggelmann, Jordan Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold · 2020
Later among the works it cites.
Evaluating the state-of-the-art of end-to-end natural language generation: The e2e nlg challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser · 2020
Later among the works it cites.
Qasc: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal · 2020
Later among the works it cites.
Explainable automated fact-checking for public health claims
Neema Kotonya and Francesca Toni · 2020
Later among the works it cites.
Data augmentation using pre-trained transformer models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho · 2020
Later among the works it cites.
Birds have four legs?! numersense: Probing numerical commonsense knowledge of pre-trained language models
Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren · 2020
Later among the works it cites.
” i’d rather just go to bed”: Understanding indirect answers
Annie Louis, Dan Roth, and Filip Radlinski · 2020
Later among the works it cites.
Limit: The literal motion in text dataset
Irene Manotas, Ngoc Phuoc An Vo, and Vadim Sheinin · 2020
Later among the works it cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee · 2020
Later among the works it cites.
Effective transfer learning for identifying similar questions: matching user questions to covid-19 faqs
Clara H McCreery, Namit Katariya, Anitha Kannan, Manish Chablani, and Xavier Amatriain · 2020
Later among the works it cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Later among the works it cites.
Biomrc: A dataset for biomedical machine reading comprehension
Dimitris Pappas, Petros Stavropoulos, Ion Androutsopoulos, and Ryan McDonald · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Getting closer to ai complete question answering: A set of prerequisite real tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky · 2020
Later among the works it cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2020
Later among the works it cites.
Investigating societal biases in a poetry composition system
Emily Sheng and David Uthus · 2020
Later among the works it cites.
Break it down: A question understanding benchmark
Tomer Wolfson, Mor Geva, Ankit Gupta, Matt Gardner, Yoav Goldberg, Daniel Deutch, and Jonathan Berant · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le · 2020
Later among the works it cites.
Muppet: Massive multi-task representations with pre-finetuning
Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, and Sonal Gupta · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf · 2021
Later among the works it cites.
Structured prediction as translation between augmented natural languages
Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cicero Nogueira dos Santos, Bing Xiang, and Stefano Soatto · 2021
Later among the works it cites.
KILT: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel · 2021
Later among the works it cites.
Meta self-training for few-shot neural sequence labeling
Yaqing Wang, Subhabrata Mukherjee, Haoda Chu, Yuancheng Tu, Ming Wu, Jing Gao, and Ahmed Hassan Awadallah · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Later among the works it cites.
CrossFit: A few-shot learning challenge for cross-task generalization in NLP
Qinyuan Ye, Bill Yuchen Lin, and Xiang Ren · 2021
Later among the works it cites.
Gpt3mix: Leveraging large-scale language models for text augmentation
Kang Min Yoo, Dongju Park, Jaewook Kang, Sang-Woo Lee, and Woomyeong Park · 2021
Later among the works it cites.
Flipda: Effective and robust data augmentation for few-shot learning
Jing Zhou, Yanan Zheng, Jie Tang, Jian Li, and Zhilin Yang · 2021
Later among the works it cites.
Ext5: Towards extreme multi-task scaling for transfer learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler · 2022
Closest in time.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M Rush · 2022
Closest in time.
PromDA: Prompt-based data augmentation for low-resource NLU tasks
Yufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, and Daxin Jiang · 2022
Closest in time.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le · 2022
Closest in time.
Zeroprompt: Scaling prompt-based pretraining to 1,000 tasks improves zero-shot generalization
Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, and Zhilin Yang · 2022
Closest in time.
FlipDA: Effective and robust data augmentation for few-shot learning
Jing Zhou, Yanan Zheng, Jie Tang, Li Jian, and Zhilin Yang · 2022
Closest in time.