Fetching the paper…
Reading the bibliography…
There have been growing concerns regarding the out-of-domain generalization ability of natural language processing (NLP) models, particularly in question-answering (QA) tasks.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 1903
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Investigating selective prediction approaches across several tasks in iid, ood, and adversarial settings
Neeraj Varshney, Swaroop Mishra, and Chitta Baral. 2022 · 2002
Earlier work this paper cites.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma. 2009 · 2009
Earlier work this paper cites.
Counterfactuals and causal inference
Stephen L Morgan and Christopher Winship. 2015 · 2015
Earlier work this paper cites.
A Semi-Supervised Learning Approach to Why-Question Answering
Jong-Hoon Oh, Kentaro Torisawa, Chikara Hashimoto, Ryu Iida, Masahiro Tanaka, and Julien Kloetzer. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
NewsQA: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. 2017 · 2017
Earlier work this paper cites.
Making neural QA as simple as possible but not simpler
Dirk Weissenborn, Georg Wiese, and Laura Seiffe. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A Smith. 2018 · 2018
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Tomas Kocisky, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gabor Melis, and Edward Grefenstette. 2018 · 2018
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F Liu, Matt Gardner, Yonatan Belinkov, Matthew E Peters, and Noah A Smith. 2019a · 2019
Earlier work this paper cites.
Unsupervised domain adaptation on reading comprehension
Yu Cao, Meng Fang, Baosheng Yu, and Joey Tianyi Zhou. 2020 · 2020
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Cited alongside, same era.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Later among the works it cites.
Template-free prompt tuning for few-shot ner
Ruotian Ma, Xin Zhou, Tao Gui, Yiding Tan, Qi Zhang, and Xuanjing Huang. 2021 · 2021
Later among the works it cites.
Shifts: A dataset of real distributional shift across multiple large-scale tasks
Andrey Malinin, Neil Band, Yarin Gal, Mark Gales, Alexander Ganshin, German Chesnokov, Alexey Noskov, Andrey Ploskonosov, Liudmila Prokhorenkova, Ivan Provilkov, Vatsal Raina, Vyas Raina, Denis Roginskiy, Mariya Shmatova, Panagiotis Tigas, and Boris Yangel. 2021 · 2021
Later among the works it cites.
Maximal multiverse learning for promoting cross-task generalization of fine-tuned language models
Itzik Malkiel and Lior Wolf. 2021 · 2021
Later among the works it cites.
Can generative pre-trained language models serve as knowledge bases for closed-book qa?
Cunxiang Wang, Pai Liu, and Yue Zhang. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What do models learn from question answering datasets?
Priyanka Sen and Amir Saffari. 2020 · 2020
Cited alongside, same era.
On the theory of transfer learning: The importance of task diversity
Nilesh Tripuraneni, Michael Jordan, and Chi Jin. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2020 · 2020
Cited alongside, same era.
Understanding and improving information transfer in multi-task learning
Sen Wu, Hongyang R Zhang, and Christopher Ré. 2020 · 2020
Cited alongside, same era.
One question answering model for many languages with cross-lingual dense passage retrieval
Akari Asai, Xinyan Yu, Jungo Kasai, and Hanna Hajishirzi. 2021 · 2021
Cited alongside, same era.
Beyond iid: three levels of generalization for question answering on knowledge bases
Yu Gu, Sue Kase, Michelle Vanni, Brian Sadler, Percy Liang, Xifeng Yan, and Yu Su. 2021 · 2021
Cited alongside, same era.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Robustness to spurious correlations in text classification via automatically generated counterfactuals
Zhao Wang and Aron Culotta. 2021 · 2021
Later among the works it cites.
Exploring the efficacy of automatically generated counterfactuals for sentiment analysis
Linyi Yang, Jiazheng Li, Pádraig Cunningham, Yue Zhang, Barry Smyth, and Ruihai Dong. 2021 · 2021
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2022 · 2022
Later among the works it cites.
A rationale-centric framework for human-in-the-loop machine learning
Jinghui Lu, Linyi Yang, Brian Mac Namee, and Yue Zhang. 2022 · 2022
Later among the works it cites.
Promda: Prompt-based data augmentation for low-resource nlu tasks
Yufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, and Daxin Jiang. 2022 · 2022
Later among the works it cites.
Towards fine-grained causal reasoning and qa
Linyi Yang, Zhen Wang, Yuxiang Wu, Jie Yang, and Yue Zhang. 2022 · 2022
Later among the works it cites.
Synthetic question value estimation for domain adaptation of question answering
Xiang Yue, Ziyu Yao, and Huan Sun. 2022 · 2022
Later among the works it cites.