Fetching the paper…
Reading the bibliography…
Recent studies on pre-trained language models have demonstrated their ability to capture factual knowledge and applications in knowledge-aware downstream tasks.
Sensebert: Driving some sense into bert
Levine, Y.; Lenz, B.; Dagan, O.; Padnos, D.; Sharir, O.; Shalev-Shwartz, S.; Shashua, A.; and Shoham, Y. 2019 · 1908
Earlier work this paper cites.
Ye, Z.-X.; Chen, Q.; Wang, W.; and Ling, Z.-H. 2019 · 1908
Earlier work this paper cites.
Informing unsupervised pretraining with external linguistic knowledge
Lauscher, A.; Vulić, I.; Ponti, E. M.; Korhonen, A.; and Glavaš, G. 2019 · 1909
Earlier work this paper cites.
Bert is not a knowledge base (yet): Factual knowledge vs. name-based reasoning in unsupervised qa
Poerner, N.; Waltinger, U.; and Schütze, H. 2019 · 1911
Earlier work this paper cites.
KEPLER: A unified model for knowledge embedding and pre-trained language representation
Wang, X.; Gao, T.; Zhu, Z.; Liu, Z.; Li, J.; and Tang, J. 2019 · 1911
Earlier work this paper cites.
oLMpics–On what Language Model Pre-training Captures
Talmor, A.; Elazar, Y.; Goldberg, Y.; and Berant, J. 2019 · 1912
Earlier work this paper cites.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?
Roberts, A.; Raffel, C.; and Shazeer, N. 2020 · 2002
Earlier work this paper cites.
K-adapter: Infusing knowledge into pre-trained models with adapters
Wang, R.; Tang, D.; Duan, N.; Wei, Z.; Huang, X.; Cao, C.; Jiang, D.; Zhou, M.; et al. 2020 · 2002
Earlier work this paper cites.
Contextualized representations using textual encyclopedic knowledge
Joshi, M.; Lee, K.; Luan, Y.; and Toutanova, K. 2020b · 2004
Earlier work this paper cites.
Lauscher, A.; Majewska, O.; Ribeiro, L. F.; Gurevych, I.; Rozanov, N.; and Glavaš, G. 2020 · 2005
Earlier work this paper cites.
Knowledge-Aware Language Model Pretraining
Rosset, C.; Xiong, C.; Phan, M.; Song, X.; Bennett, P.; and Tiwary, S. 2020 · 2007
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Cited alongside, same era.
WiNER: A Wikipedia annotated corpus for named entity recognition
Ghaddar, A.; and Langlais, P. 2017 · 2017
Cited alongside, same era.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Joshi, M.; Choi, E.; Weld, D. S.; and Zettlemoyer, L. 2017 · 2017
Cited alongside, same era.
LIMIT-BERT: Linguistic informed multi-task bert
Zhou, J.; Zhang, Z.; and Zhao, H. 2019 · 2017
Cited alongside, same era.
MRQA 2019 Shared Task: Evaluating Generalization in Reading Comprehension
Fisch, A.; Talmor, A.; Jia, R.; Seo, M.; Choi, E.; and Chen, D. 2019 · 2019
Later among the works it cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 2019
Later among the works it cites.
Barack’s Wife Hillary: Using Knowledge Graphs for Fact-Aware Language Modeling
Logan, R.; Liu, N. F.; Peters, M. E.; Gardner, M.; and Singh, S. 2019 · 2019
Later among the works it cites.
Knowledge Enhanced Contextual Word Representations
Peters, M. E.; Neumann, M.; Logan, R.; Schwartz, R.; Joshi, V.; Singh, S.; and Smith, N. A. 2019 · 2019
Later among the works it cites.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Lewis, P.; Bakhtin, A.; Wu, Y.; Miller, A. H.; and Riedel, S. 2019 · 2019
Later among the works it cites.
ERNIE: Enhanced Language Representation with Informative Entities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J. G.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Joshi, M.; Chen, D.; Liu, Y.; Weld, D. S.; Zettlemoyer, L.; and Levy, O. 2020a
Cited in the paper.
Zhang, Z.; Han, X.; Liu, Z.; Jiang, X.; Sun, M.; and Liu, Q. 2019 · 2019
Later among the works it cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Clark, K.; Luong, M.-T.; Le, Q. V.; and Manning, C. D. 2020 · 2020
Closest in time.
Ernie 2.0: A continual pre-training framework for language understanding
Sun, Y.; Wang, S.; Li, Y.; Feng, S.; Tian, H.; Wu, H.; and Wang, H. 2020 · 2020
Closest in time.
Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model
Xiong, W.; Du, J.; Wang, W. Y.; and Stoyanov, V. 2020 · 2020
Closest in time.