Fetching the paper…
Reading the bibliography…
Standard test sets for supervised learning evaluate in-distribution generalization.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Alane Suhr and Yoav Artzi. 2019 · 1909
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2019 · 1910
Earlier work this paper cites.
BLiMP: A benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2019 · 1912
Earlier work this paper cites.
Phonemics: A Technique for Reducing Languages to Writing
K.L. Pike. 1946 · 1946
Earlier work this paper cites.
Building a large annotated corpus of english: The Penn treebank
Mitchell Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Prepositional phrase attachment through a backed-off model
Michael Collins and James Brooks. 1995 · 1995
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira. 2000 · 2000
Earlier work this paper cites.
LinES: an English-Swedish parallel treebank
Lars Ahrenberg. 2007 · 2007
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan. 2010 · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2011 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
SemEval-2013 Task 1: TempEval-3: Evaluating time expressions, events, and temporal relations
Naushad UzZaman, Hector Llorens, Leon Derczynski, James Allen, Marc Verhagen, and James Pustejovsky. 2013 · 2013
Earlier work this paper cites.
A gold standard dependency corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel Bowman, Miriam Connor, John Bauer, and Chris Manning. 2014 · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
PartTUT: The Turin university parallel treebank
Manuela Sanguinetti and Cristina Bosco. 2015 · 2015
Earlier work this paper cites.
A thorough examination of the CNN/Daily Mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
Universal dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajič, Christopher D. Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman. 2016 · 2016
Earlier work this paper cites.
Evaluating the morphological competence of machine translation systems
Franck Burlot and François Yvon. 2017 · 2017
Earlier work this paper cites.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
Towards linguistically generalizable NLP systems: A workshop and shared task
Allyson Ettinger, Sudha Rao, Hal Daumé III, and Emily M. Bender. 2017 · 2017
Earlier work this paper cites.
A challenge set approach to evaluating machine translation
Pierre Isabelle, Colin Cherry, and George Foster. 2017 · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
RACE: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Cited alongside, same era.
How grammatical is character-level neural machine translation? Assessing MT quality with contrastive translation pairs
Rico Sennrich. 2017 · 2017
Cited alongside, same era.
Foil it! Find One mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017 · 2017
Cited alongside, same era.
The GUM corpus: Creating multilayer resources in the classroom
Amir Zeldes. 2017 · 2017
Cited alongside, same era.
The WMT’18 morpheval test suites for English-Czech, English-German, English-Finnish and Turkish-English
Franck Burlot, Yves Scherrer, Vinit Ravishankar, Ondřej Bojar, Stig-Arne Grönroos, Maarit Koponen, Tommi Nieminen, and François Yvon. 2018 · 2018
Cited alongside, same era.
Misleading failures of partial-input baselines
Shi Feng, Eric Wallace, and Jordan Boyd-Graber. 2019 · 2019
Later among the works it cites.
Are we modeling the task or the annotator? An investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
A multi-type multi-span network for reading comprehension that requires discrete reasoning
Minghao Hu, Yuxing Peng, Zhen Huang, and Dongsheng Li. 2019 · 2019
Later among the works it cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019 · 2019
Later among the works it cites.
Reasoning over paragraph effects in situations
Kevin Lin, Oyvind Tafjord, Peter Clark, and Matt Gardner. 2019 · 2019
Later among the works it cites.
On evaluation of adversarial perturbations for sequence-to-sequence models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Cited alongside, same era.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Detecting and correcting for label shift with black box predictors
Zachary Lipton, Yu-Xiang Wang, and Alexander Smola. 2018 · 2018
Cited alongside, same era.
Gender bias in neural natural language processing
Kaiji Lu, Piotr Mardziel, Fangjing Wu, Preetam Amancharla, and Anupam Datta. 2018 · 2018
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen. 2018 · 2018
Cited alongside, same era.
Paul Michel, Xian Li, Graham Neubig, and Juan Miguel Pino. 2019 · 2019
Later among the works it cites.
Compositional questions do not necessitate multi-hop reasoning
Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
An Improved Neural Baseline for Temporal Relation Extraction
Qiang Ning, Sanjay Subramanian, and Dan Roth. 2019 · 2019
Later among the works it cites.
Probing neural network comprehension of natural language arguments
Timothy Niven and Hung-Yu Kao. 2019 · 2019
Later among the works it cites.
Do ImageNet classifiers generalize to ImageNet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019 · 2019
Later among the works it cites.
Are red roses red? Evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Later among the works it cites.
Learning a SAT solver from single-bit supervision
Daniel Selsam, Matthew Lamm, Benedikt Bünz, Percy Liang, Leonardo de Moura, and David L. Dill. 2019 · 2019
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
QuaRTz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019 · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2019 · 2019
Later among the works it cites.
Cold case: The lost MNIST digits
Chhavi Yadav and Léon Bottou. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
“Going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth. 2019 · 2019
Later among the works it cites.
Good-enough compositional data augmentation
Jacob Andreas. 2020 · 2020
Closest in time.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C Lipton. 2020 · 2020
Closest in time.
More bang for your buck: Natural perturbation for robust question answering
Daniel Khashabi, Tushar Khot, and Ashish Sabhwawal. 2020 · 2020
Closest in time.
Predictive biases in natural language processing models: A conceptual framework and overview
Deven Shah, H. Andrew Schwartz, and Dirk Hovy. 2020 · 2020
Closest in time.