Fetching the paper…
Reading the bibliography…
Despite data's crucial role in machine learning, most existing tools and research tend to focus on systems on top of existing data rather than how to interpret and manipulate data.
Type/token ratios: what do they really tell us?
Brian Richards. 1987 · 1987
Earlier work this paper cites.
Term-weighting approaches in automatic text retrieval
Gerard Salton and Christopher Buckley. 1988 · 1988
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Destructive messages: How hate speech paves the way for harmful social movements , volume 27
Alexander Tsesis. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Multi-dimensional gender bias classification
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, and Adina Williams. 2020 · 2005
Earlier work this paper cites.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma. 2009 · 2009
Earlier work this paper cites.
An NLP curator (or: How I learned to stop worrying and love NLP pipelines)
James Clarke, Vivek Srikumar, Mark Sammons, and Dan Roth. 2012 · 2012
Earlier work this paper cites.
From amateurs to connoisseurs: modeling the evolution of user expertise through online reviews
Julian John McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
Christopher D Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Deep learning for hate speech detection in tweets
Pinkesh Badjatiya, Shashank Gupta, Manish Gupta, and Vasudeva Varma. 2017 · 2017
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael W. Macy, and Ingmar Weber. 2017 · 2017
Earlier work this paper cites.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Matthew Honnibal and Ines Montani. 2017 · 2017
Earlier work this paper cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Earlier work this paper cites.
AllenNLP: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018a · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018b · 2018
Cited alongside, same era.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Later among the works it cites.
Sentence piece
Taku Kudo. 2020 · 2020
Later among the works it cites.
Snorkel: rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason A. Fries, Sen Wu, and Christopher Ré. 2020 · 2020
Later among the works it cites.
Systematic inequalities in language technology performance across the world’s languages
Damián Blasi, Antonios Anastasopoulos, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Nl-augmenter: A framework for task-sensitive natural language augmentation
Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, et al. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass. 2019 · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
EDA: easy data augmentation techniques for boosting performance on text classification tasks
Jason W. Wei and Kai Zou. 2019 · 2019
Cited alongside, same era.
Open domain web keyphrase extraction beyond language modeling
Lee Xiong, Chuan Hu, Chenyan Xiong, Daniel Fernando Campos, and Arnold Overwijk. 2019 · 2019
Cited alongside, same era.
A closer look at data bias in neural extractive summarization models
Ming Zhong, Danqing Wang, Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2019 · 2019
Cited alongside, same era.
Dataset geography: Mapping language data to language users
Fahim Faisal, Yinkai Wang, and Antonios Anastasopoulos. 2021 · 2021
Later among the works it cites.
A survey of data augmentation approaches for NLP
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard H. Hovy. 2021 · 2021
Later among the works it cites.
Explore datasets in know your data
Google. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Sasko, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander M. Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Later among the works it cites.
Data and its (dis) contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. 2021 · 2021
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with checklist (extended abstract)
Marco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2021 · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021 · 2021
Later among the works it cites.
How well do you know your summarization datasets?
Priyam Tejaswin, Dhruv Naik, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Tensorflow datasets, a collection of ready-to-use datasets
TFData. 2021 · 2021
Later among the works it cites.
TextFlint: Unified multilingual robustness evaluation toolkit for natural language processing
Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, Qinzhuo Wu, Zhengyan Li, Chong Zhang, Ruotian Ma, Zichu Fei, Ruijian Cai, Jun Zhao, Xingwu Hu, Zhiheng Yan, Yiding Tan, Yuan Hu, Qiyuan Bian, Zhihua Liu, Shan Qin, Bolin Zhu, Xiaoyu Xing, Jinlan Fu, Yue Zhang, Minlong Peng, Xiaoqing Zheng, Yaqian Zhou, Zhongyu Wei, Xipeng Qiu, and Xuanjing Huang. 2021 · 2021
Later among the works it cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.