Fetching the paper…
Reading the bibliography…
We study the robustness of machine reading comprehension (MRC) models to entity renaming -- do models make more wrong predictions when the same questions are asked about an entity whose name has been changed? Such failures imply that models overly rely on entity information to answer questions, and thus may generalize poorly when facts about the world change or questions are asked about novel entities.
UNIFIEDQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
Oshin Agarwal, Yinfei Yang, Byron C. Wallace, and A. Nenkova. 2020 · 2004
Earlier work this paper cites.
Overview of qa4mre at clef 2011: Question answering for machine reading evaluation
Anselmo Peñas, Eduard H Hovy, Pamela Forner, Álvaro Rodrigo, Richard Sutcliffe, Corina Forascu, and Caroline Sporleder. 2011 · 2011
Earlier work this paper cites.
A thorough examination of the CNN/Daily Mail reading comprehension task
Danqi Chen, Jason Bolton, and Christopher D. Manning. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Searchqa: A new q&a dataset augmented with context from a search engine
Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C. Lipton. 2018 · 2018
Earlier work this paper cites.
Semantically equivalent adversarial rules for debugging NLP models
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018 · 2018
Earlier work this paper cites.
Are you tough enough? framework for robustness validation of machine comprehension systems
Barbara Rychalska, Dominika Basaj, and Przemyslaw Biecek. 2018 · 2018
Earlier work this paper cites.
What makes reading comprehension questions easier?
Saku Sugawara, Kentaro Inui, Satoshi Sekine, and Akiko Aizawa. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
MRQA 2019 shared task: Evaluating generalization in reading comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019 · 2019
Cited alongside, same era.
Improving the robustness of question answering systems to question paraphrasing
Wee Chung Gan and Hwee Tou Ng. 2019 · 2019
Cited alongside, same era.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019 · 2019
Cited alongside, same era.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Cited alongside, same era.
Syntactic data augmentation increases robustness to inference heuristics
Junghyun Min, R. Thomas McCoy, Dipanjan Das, Emily Pitler, and Tal Linzen. 2020 · 2020
Later among the works it cites.
Enhancing natural language inference using new and expanded training data sets and new learning models
Arindam Mitra, Ishan Shrivastava, and Chitta Baral. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
What do models learn from question answering datasets?
Priyanka Sen and Amir Saffari. 2020 · 2020
Later among the works it cites.
“you are grounded!”: Latent name artifacts in pre-trained language models
Vered Shwartz, Rachel Rudinger, and Oyvind Tafjord. 2020 · 2020
Later among the works it cites.
Assessing the benchmarking capacity of machine reading comprehension datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Combating adversarial misspellings with robust word recognition
Danish Pruthi, Bhuwan Dhingra, and Zachary C. Lipton. 2019 · 2019
Cited alongside, same era.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Improving generalization in coreference resolution via adversarial training
Sanjay Subramanian and Dan Roth. 2019 · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
Co-attention hierarchical network: Generating coherent long distractors for reading comprehension
Xiaorui Zhou, Senlin Luo, and Yunfang Wu. 2020 · 2019
Cited alongside, same era.
What’s in a name? are BERT named entity representations just as good for any other name?
Sriram Balasubramanian, Naman Jain, Gaurav Jindal, Abhijeet Awasthi, and Sunita Sarawagi. 2020 · 2020
Cited alongside, same era.
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020 · 2020
Later among the works it cites.
Checkdst: Measuring real-world generalization of dialogue state tracking performance
Hyundong Cho, Chinnadhurai Sankar, Christopher Lin, Kaushik Ram Sadagopan, Shahin Shayandeh, Asli Celikyilmaz, Jonathan May, and Ahmad Beirami. 2021 · 2021
Closest in time.
Why machine reading comprehension models learn shortcuts?
Yuxuan Lai, Chen Zhang, Yansong Feng, Quzhe Huang, and Dongyan Zhao. 2021 · 2021
Closest in time.
Bill Yuchen Lin, Wenyang Gao, Jun Yan, Ryan Moreno, and Xiang Ren. 2021 · 2021
Closest in time.
Challenges in generalization in open domain question answering
Linqing Liu, Patrick Lewis, Sebastian Riedel, and Pontus Stenetorp. 2021 · 2021
Closest in time.
Entity-based knowledge conflicts in question answering
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021 · 2021
Closest in time.
Accuracy on the line: On the strong correlation between out-of-distribution and in-distribution generalization
John P Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt. 2021 · 2021
Closest in time.
Benchmarking robustness of machine reading comprehension models
Chenglei Si, Ziqing Yang, Yiming Cui, Wentao Ma, Ting Liu, and Shijin Wang. 2021 · 2021
Closest in time.
On the influence of masking policies in intermediate pre-training
Qinyuan Ye, Belinda Z Li, Sinong Wang, Benjamin Bolte, Hao Ma, Wen-tau Yih, Xiang Ren, and Madian Khabsa. 2021 · 2021
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.