Fetching the paper…
Reading the bibliography…
Recent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines (e.g., hypothesis-only models for SNLI or question-only models for VQA).
Test-Driven Development by Example
Kent Beck. 2002 · 2002
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
A thorough examination of the CNN/Daily Mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Blindfold baselines for embodied QA
Ankesh Anand, Eugene Belilovsky, Kyle Kastner, Hugo Larochelle, and Aaron Courville. 2018 · 2018
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
Breaking NLI systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke S. Zettlemoyer. 2018 · 2018
Cited alongside, same era.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton. 2018 · 2018
Cited alongside, same era.
Visual dialogue without vision or dialogue
Daniela Massiceti, Puneet K. Dokania, N. Siddharth, and Philip H.S. Torr. 2018 · 2018
Cited alongside, same era.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Cited alongside, same era.
Deep k-nearest neighbors: Towards confident, interpretable and robust deep learning
Nicolas Papernot and Patrick D. McDaniel. 2018 · 2018
Cited alongside, same era.
Hypothesis only baselines in natural language inference
Please stop explaining black box models for high stakes decisions
Cynthia Rudin. 2018 · 2018
Later among the works it cites.
Interpreting neural networks with nearest neighbors
Eric Wallace, Shi Feng, and Jordan Boyd-Graber. 2018 · 2018
Later among the works it cites.
Provable defenses against adversarial examples via the convex outer adversarial polytope
Eric Wong and J. Zico Kolter. 2018 · 2018
Later among the works it cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Y. Zou. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Certified defenses against adversarial examples
Aditi Raghunathan, Jacob Steinhardt, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Semantically equivalent adversarial rules for debugging NLP models
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Cited alongside, same era.
Closest in time.
Shifting the baseline: Single modality performance on visual navigation & QA
Jesse Thomason, Daniel Gordan, and Yonatan Bisk. 2019 · 2019
Closest in time.
Trick me if you can: Human-in-the-loop generation of adversarial examples for question answering
Eric Wallace, Pedro Rodriguez, Shi Feng, Ikuya Yamada, and Jordan Boyd-Graber. 2019 · 2019
Closest in time.