Fetching the paper…
Reading the bibliography…
Large amounts of training data are one of the major reasons for the high performance of state-of-the-art NLP models.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Smoking and lung cancer: recent evidence and a discussion of some questions
Jerome Cornfield, William Haenszel, E Cuyler Hammond, Abraham M Lilienfeld, Michael B Shimkin, and Ernst L Wynder. 1959 · 1959
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R Hampel. 1974 · 1974
Earlier work this paper cites.
Probabilistic reasoning in intelligent systems: networks of plausible inference
Judea Pearl. 1988 · 1988
Earlier work this paper cites.
Sensitivity analysis for selection bias and unmeasured confounding in missing data and causal inference models
James M Robins, Andrea Rotnitzky, and Daniel O Scharfstein. 2000 · 2000
Earlier work this paper cites.
Causality
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Matching methods for causal inference: A review and a look forward
Elizabeth A Stuart. 2010 · 2010
Earlier work this paper cites.
Sensitivity analysis for causal inference under unmeasured confounding and measurement error problems
Iván Díaz and Mark J van der Laan. 2013 · 2013
Earlier work this paper cites.
An introduction to sensitivity analysis for unobserved confounding in nonexperimental prevention research
Weiwei Liu, S Janet Kuramoto, and Elizabeth A Stuart. 2013 · 2013
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
T-rex: A large scale alignment of natural language with knowledge base triples
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018 · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A Smith. 2018 · 2018
Earlier work this paper cites.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary C Lipton. 2018 · 2018
Earlier work this paper cites.
Stress test evaluation for natural language inference
Aakanksha Naik, Abhilasha Ravichander, Norman Sadeh, Carolyn Rose, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Performance impact caused by hidden bias of training data for recognizing textual entailment
Masatoshi Tsuchiya. 2018 · 2018
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
The emergence of number and syntax units in lstm language models
Yair Lakretz, Germán Kruszewski, Théo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni. 2019 · 2019
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
An analysis of dataset overlap on Winograd-style tasks
Ali Emami, Kaheer Suleman, Adam Trischler, and Jackie Chi Kit Cheung. 2020 · 2020
Cited alongside, same era.
Static embeddings as efficient knowledge bases?
Philipp Dufter, Nora Kassner, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Back to square one: Artifact detection, training and commonsense disentanglement in the winograd schema
Yanai Elazar, Hongming Zhang, Yoav Goldberg, and Dan Roth. 2021c · 2021
Later among the works it cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021 · 2021
Later among the works it cites.
Causal direction of data collection matters: Implications of causal and anticausal learning for nlp
Zhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, and Bernhard Schoelkopf. 2021 · 2021
Later among the works it cites.
BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief
Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark. 2021b · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Allyson Ettinger. 2020 · 2020
Cited alongside, same era.
X-FACTR: Multilingual factual knowledge retrieval from pretrained language models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Are pretrained language models symbolic reasoners over knowledge?
Nora Kassner, Benno Krojer, and Hinrich Schütze. 2020 · 2020
Cited alongside, same era.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Nora Kassner and Hinrich Schütze. 2020 · 2020
Cited alongside, same era.
How context affects language models’ factual predictions
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2020 · 2020
Cited alongside, same era.
E-BERT: Efficient-yet-effective entity embeddings for BERT
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
R Thomas McCoy, Paul Smolensky, Tal Linzen, Jianfeng Gao, and Asli Celikyilmaz. 2021 · 2021
Later among the works it cites.
Counterfactual interventions reveal the causal effect of relative clause representations on agreement prediction
Shauli Ravfogel, Grusha Prasad, Tal Linzen, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Language models as or for knowledge bases
Simon Razniewski, Andrew Yates, Nora Kassner, and Gerhard Weikum. 2021 · 2021
Later among the works it cites.
The multiberts: Bert reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al. 2021 · 2021
Later among the works it cites.
Mediators in determining what processing bert performs first
Aviv Slobodkin, Leshem Choshen, and Omri Abend. 2021 · 2021
Later among the works it cites.
Counterfactual invariance to spurious correlations: Why and how to pass stress tests
Victor Veitch, Alexander D’Amour, Steve Yadlowsky, and Jacob Eisenstein. 2021 · 2021
Later among the works it cites.
Frequency effects on syntactic rule learning in transformers
Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick. 2021 · 2021
Later among the works it cites.
Causal distillation for language models
Zhengxuan Wu, Atticus Geiger, Josh Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, and Noah D Goodman. 2021 · 2021
Later among the works it cites.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2021 · 2021
Later among the works it cites.
Tracing knowledge in language models back to the training data
Ekin Akyürek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu. 2022 · 2022
Closest in time.
On the pitfalls of analyzing individual neurons in language models
Omer Antverg and Yonatan Belinkov. 2022 · 2022
Closest in time.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Closest in time.
Informativeness and invariance: Two perspectives on spurious correlations in natural language
Jacob Eisenstein. 2022 · 2022
Closest in time.
Impact of pretraining term frequencies on few-shot reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022 · 2022
Closest in time.