Fetching the paper…
Reading the bibliography…
Previous literatures show that pre-trained masked language models (MLMs) such as BERT can achieve competitive factual knowledge extraction performance on some datasets, indicating that MLMs can potentially be a reliable knowledge source.
Assessing BERT’s Syntactic Abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Do Neural Language Representations Learn Physical Commonsense?
Maxwell Forbes, Ari Holtzman, and Yejin Choi. 2019 · 1908
Earlier work this paper cites.
Do Attention Heads in BERT Track Syntactic Dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R. Bowman. 2019 · 1911
Earlier work this paper cites.
A taxonomy of problems with fast parallel algorithms
Stephen A. Cook. 1985 · 1985
Earlier work this paper cites.
Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems
Andrea Madotto, Zihan Liu, Zhaojiang Lin, and Pascale Fung. 2020 · 2008
Earlier work this paper cites.
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Timo Schick and Hinrich Schütze. 2020b · 2009
Earlier work this paper cites.
Language Models are Open Knowledge Graphs
Chenguang Wang, Xiao Liu, and Dawn Song. 2020 · 2010
Earlier work this paper cites.
Making Pre-trained Language Models Better Few-shot Learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2020 · 2012
Earlier work this paper cites.
Few-Shot Text Generation with Pattern-Exploiting Training
Timo Schick and Hinrich Schütze. 2020a · 2012
Earlier work this paper cites.
A neural network for factoid question answering over paragraphs
Mohit Iyyer, Jordan Boyd-Graber, Leonardo Claudino, Richard Socher, and Hal Daumé III. 2014 · 2014
Earlier work this paper cites.
Wikidata: A free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Evaluating Commonsense in Pre-trained Language Models
Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang. 2019 · 2014
Earlier work this paper cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Earlier work this paper cites.
Reading Wikipedia to answer open-domain questions
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Earlier work this paper cites.
T-REx: A large scale alignment of natural language with knowledge base triples
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Commonsense knowledge mining from pretrained models
Joe Davison, Joshua Feldman, and Alexander Rush. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Open sesame: Getting inside BERT’s linguistic knowledge
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Later among the works it cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Later among the works it cites.
X-FACTR: Multilingual factual knowledge retrieval from pretrained language models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020a · 2020
Later among the works it cites.
Are pretrained language models symbolic reasoners over knowledge?
Nora Kassner, Benno Krojer, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
BERT-kNN: Adding a kNN search component to pretrained language models for better QA
Nora Kassner and Hinrich Schütze. 2020a · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019a · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Birds have four legs?! NumerSense: Probing Numerical Commonsense Knowledge of Pre-Trained Language Models
Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. 2020 · 2020
Later among the works it cites.
How context affects language models’ factual predictions
Fabio Petroni, Patrick S. H. Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020 · 2020
Later among the works it cites.
E-BERT: Efficient-yet-effective entity embeddings for BERT
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Later among the works it cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
“you are grounded!”: Latent name artifacts in pre-trained language models
Vered Shwartz, Rachel Rudinger, and Oyvind Tafjord. 2020 · 2020
Later among the works it cites.
Pre-training is (almost) all you need: An application to commonsense reasoning
Alexandre Tamborrino, Nicola Pellicanò, Baptiste Pannier, Pascal Voitot, and Louise Naudin. 2020 · 2020
Later among the works it cites.
Benchmarking Knowledge-Enhanced Commonsense Question Answering via Knowledge-to-Text Transformation
Ning Bian, Xianpei Han, Bo Chen, and Le Sun. 2021 · 2021
Closest in time.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Closest in time.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2021 · 2021
Closest in time.