Fetching the paper…
Reading the bibliography…
In order to reliably process natural language, NLP systems must generalize to the long tail of rare utterances.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A model and an hypothesis for language structure
Victor H. Yngve. 1960 · 1960
Earlier work this paper cites.
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
J. Peter Kincaid, Robert P. Fishburne, R L Rogers, and Brad S. Chissom. 1975 · 1975
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
J. Fodor and Z. Pylyshyn. 1988 · 1988
Earlier work this paper cites.
Measures of distance between probability distributions
J.K Chung, P.L Kannappan, C.T Ng, and P.K Sahoo. 1989 · 1989
Earlier work this paper cites.
Nltk: The natural language toolkit
Edward Loper and Steven Bird. 2002 · 2002
Earlier work this paper cites.
Syntactic complexity measures for detecting mild cognitive impairment
Brian Roark, Margaret Mitchell, and Kristy Hollingshead. 2007 · 2007
Earlier work this paper cites.
Moving beyond Kucera and Francis: a critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English
Marc Brysbaert and Boris New. 2009 · 2009
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Improving text-to-SQL evaluation methodology
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev. 2018 · 2018
Earlier work this paper cites.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein. 2018 · 2018
Earlier work this paper cites.
Brenden M. Lake and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. 2018 · 2018
Earlier work this paper cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Earlier work this paper cites.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018 · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Don’t forget the long tail! a comprehensive analysis of morphological generalization in bilingual lexicon induction
Paula Czarnowska, Sebastian Ruder, Edouard Grave, Ryan Cotterell, and Ann Copestake. 2019 · 2019
Earlier work this paper cites.
MRQA 2019 shared task: Evaluating generalization in reading comprehension
Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. 2019 · 2019
Earlier work this paper cites.
Posing fair generalization tasks for natural language inference
Atticus Geiger, Ignacio Cases, Lauri Karttunen, and Christopher Potts. 2019 · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Likelihood ratios for out-of-distribution detection
Jie Ren, Peter J. Liu, Emily Fertig, Jasper Snoek, Ryan Poplin, Mark Depristo, Joshua Dillon, and Balaji Lakshminarayanan. 2019 · 2019
Cited alongside, same era.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Utility is in the eye of the user: A critique of NLP leaderboards
Kawin Ethayarajh and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Cited alongside, same era.
Competency problems: On finding and removing artifacts in language data
Matt Gardner, William Merrill, Jesse Dodge, Matthew Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Beyond i.i.d.: Three levels of generalization for question answering on knowledge bases
Yu Gu, Sue Kase, Michelle Vanni, Brian Sadler, Percy Liang, Xifeng Yan, and Yu Su. 2021 · 2021
Later among the works it cites.
On the efficacy of adversarial data collection for question answering: Results from a large-scale randomized study
Divyansh Kaushik, Douwe Kiela, Zachary C. Lipton, and Wen-tau Yih. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pretrained transformers improve out-of-distribution robustness
Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song. 2020 · 2020
Cited alongside, same era.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
Cited alongside, same era.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Sch"arli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang, Marc van Zee, and Olivier Bousquet. 2020 · 2020
Cited alongside, same era.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen. 2020 · 2020
Cited alongside, same era.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew E. Peters, Ashish Sabharwal, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
How can we accelerate progress towards human-like linguistic generalization?
Tal Linzen. 2020 · 2020
Cited alongside, same era.
The effect of natural distribution shift on question answering models
John Miller, Karl Krauth, Benjamin Recht, and Ludwig Schmidt. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Challenges in generalization in open domain question answering
Linqing Liu, Patrick S. H. Lewis, Sebastian Riedel, and Pontus Stenetorp. 2021 · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
John P Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
Bootleg: Chasing the tail with self-supervised named entity disambiguation
Laurel J. Orr, Megan Leszczynski, Neel Guha, Sen Wu, Simran Arora, Xiao Ling, and Christopher Ré. 2021 · 2021
Later among the works it cites.
Adversarially constructed evaluation sets are more challenging, but may not be fair
Jason Phang, Angelica Chen, William Huang, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Revisiting few-shot relation classification: Evaluation data and classification schemes
Ofer Sabo, Yanai Elazar, Yoav Goldberg, and Ido Dagan. 2021 · 2021
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Simple entity-centric questions challenge dense retrievers
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Compositional generalization and natural language variation: Can a semantic parsing approach handle both?
Peter Shaw, Ming-Wei Chang, Panupong Pasupat, and Kristina Toutanova. 2021 · 2021
Later among the works it cites.
We need to talk about random splits
Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, and Katja Filippova. 2021 · 2021
Later among the works it cites.
Analyzing dynamic adversarial training data in the limit
Eric Wallace, Adina Williams, Robin Jia, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Measuring causal effects of data statistics on language model’s ‘factual’ predictions
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Amir Feder, Abhilasha Ravichander, Marius Mosbach, Yonatan Belinkov, Hinrich Schütze, and Yoav Goldberg. 2022 · 2022
Closest in time.
Large language models struggle to learn long-tail knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022 · 2022
Closest in time.
Adapting to the long tail: A meta-analysis of transfer learning research for language understanding tasks
Aakanksha Naik, Jill Lehman, and Carolyn Rosé. 2022 · 2022
Closest in time.
Impact of pretraining term frequencies on few-shot reasoning
Yasaman Razeghi, Robert L. Logan, Matt Gardner, and Sameer Singh. 2022 · 2022
Closest in time.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022 · 2022
Closest in time.