Fetching the paper…
Reading the bibliography…
Developing methods to adversarially challenge NLP systems is a promising avenue for improving both model performance and interpretability.
On comprehending sentences: Syntactic parsing strategies
Lynn Frazier. 1979 · 1979
Earlier work this paper cites.
Easy victories and uphill battles in coreference resolution
Greg Durrett and Dan Klein. 2013 · 1982
Earlier work this paper cites.
Making and correcting errors during sentence comprehension: Eye movements in the analysis of structurally ambiguous sentences
Lyn Frazier and Keith Rayner. 1982 · 1982
Earlier work this paper cites.
Relevance: Communication and cognition , volume 142
Dan Sperber and Deirdre Wilson. 1986 · 1986
Earlier work this paper cites.
Recovery from misanalyses of garden-path sentences
Fernanda Ferreira and John M Henderson. 1991 · 1991
Earlier work this paper cites.
The child’s theory of mind
Henry M Wellman. 1992 · 1992
Earlier work this paper cites.
Using language
Herbert H Clark. 1996 · 1996
Earlier work this paper cites.
Linguistic complexity: Locality of syntactic dependencies
Edward Gibson. 1998 · 1998
Earlier work this paper cites.
Beat the AI: investigating adversarial human annotations for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020 · 2002
Earlier work this paper cites.
The next decade in ai: four steps towards robust artificial intelligence
Gary Marcus. 2020 · 2002
Earlier work this paper cites.
Look at the first sentence: Position bias in question answering
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang. 2020 · 2004
Earlier work this paper cites.
Expectation-based syntactic comprehension
Roger Levy. 2008 · 2008
Earlier work this paper cites.
How language production shapes language form and comprehension
Maryellen C. MacDonald. 2013 · 2013
Earlier work this paper cites.
A functional dissociation between language and multiple-demand systems revealed in patterns of bold signal fluctuations
Idan Blank, Nancy Kanwisher, and Evelina Fedorenko. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Large-scale evidence of dependency length minimization in 37 languages
Richard Futrell, Kyle Mahowald, and Edward Gibson. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Cited alongside, same era.
ETPC - a paraphrase identification corpus annotated with extended paraphrase typology and negation
Venelin Kovatchev, M. Antònia Martí, and Maria Salamó. 2018 · 2018
Cited alongside, same era.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition
Paloma Jeretic, Alex Warstadt, Suvrat Bhooshan, and Adina Williams. 2020 · 2020
Later among the works it cites.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Later among the works it cites.
“what is on your mind?” automated scoring of mindreading in childhood and early adolescence
Venelin Kovatchev, Phillip Smith, Mark Lee, Imogen Grumley Traynor, Irene Luque Aguilera, and Rory Devine. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
ConjNLI: Natural language inference over conjunctive sentences
Swarnadeep Saha, Yixin Nie, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Cited alongside, same era.
How efficiency shapes human language
Edward Gibson, Richard Futrell, Steven P Piantadosi, Isabelle Dautriche, Kyle Mahowald, Leon Bergen, and Roger Levy. 2019 · 2019
Cited alongside, same era.
Reference resolution
Elsi Kaiser and Emily Fedele. 2019 · 2019
Cited alongside, same era.
A qualitative evaluation framework for paraphrase identification
Venelin Kovatchev, M. Antonia Marti, Maria Salamo, and Javier Beltran. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Analyzing compositionality-sensitivity of nli models
Yixin Nie, Yicheng Wang, and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
Trick me if you can: Human-in-the-loop generation of adversarial examples for question answering
Eric Wallace, Pedro Rodriguez, Shi Feng, Ikuya Yamada, and Jordan Boyd-Graber. 2019 · 2019
Cited alongside, same era.
olmpics-on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Later among the works it cites.
Back to square one: Artifact detection, training and commonsense disentanglement in the Winograd schema
Yanai Elazar, Hongming Zhang, Yoav Goldberg, and Dan Roth. 2021 · 2021
Later among the works it cites.
On the efficacy of adversarial data collection for question answering: Results from a large-scale randomized study
Divyansh Kaushik, Douwe Kiela, Zachary C. Lipton, and Wen-tau Yih. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
Adversarially constructed evaluation sets are more challenging, but may not be fair
Jason Phang, Angelica Chen, William Huang, and Samuel R Bowman. 2021 · 2021
Later among the works it cites.
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Learning from the worst: Dynamically generated datasets to improve online hate detection
Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Frequency effects on syntactic rule learning in transformers
Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021 · 2021
Later among the works it cites.
The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail
Samuel Bowman. 2022 · 2022
Closest in time.
ANLIzing the adversarial natural language inference dataset
Adina Williams, Tristan Thrush, and Douwe Kiela. 2022 · 2022
Closest in time.