Fetching the paper…
Reading the bibliography…
We present a novel way of injecting factual knowledge about entities into the pretrained BERT model (Devlin et al., 2019): We align Wikipedia2Vec entity vectors (Yamada et al., 2016) with BERT's native wordpiece vector space and use the aligned entity vectors as if they were wordpiece vectors.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
StructBERT: Incorporating language structures into pre-training for deep language understanding
Wei Wang, Bin Bi, Ming Yan, Chen Wu, Zuyi Bao, Liwei Peng, and Luo Si. 2019b · 1908
Earlier work this paper cites.
BERTRAM: Improved word embeddings have big impact on contextualized model performance
Timo Schick and Hinrich Schütze. 2019 · 1910
Earlier work this paper cites.
YELM: End-to-end contextualized entity linking
Haotian Chen, Sahil Wadhwa, Xi David Li, and Andrej Zukov-Gregoric. 2019 · 1911
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2019 · 1911
Earlier work this paper cites.
KEPLER: A unified model for knowledge embedding and pre-trained language representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2019c · 1911
Earlier work this paper cites.
K-Adapter: Infusing knowledge into pre-trained models with adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Cuihong Cao, Daxin Jiang, Ming Zhou, et al. 2020 · 2002
Earlier work this paper cites.
Introduction to Information Retrieval
Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008 · 2008
Earlier work this paper cites.
Robust disambiguation of named entities in text
Johannes Hoffart, Mohamed Amir Yosef, Ilaria Bordino, Hagen Fürstenau, Manfred Pinkal, Marc Spaniol, Bilyana Taneva, Stefan Thater, and Gerhard Weikum. 2011 · 2011
Earlier work this paper cites.
Joint learning of the embedding of words and entities for named entity disambiguation
Ikuya Yamada, Hiroyuki Shindo, Hideaki Takeda, and Yoshiyasu Takefuji. 2016 · 2016
Cited alongside, same era.
Question answering on knowledge bases and text using universal schema and memory networks
Rajarshi Das, Manzil Zaheer, Siva Reddy, and Andrew McCallum. 2017 · 2017
Cited alongside, same era.
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Samuel L Smith, David HP Turban, Steven Hamblin, and Nils Y Hammerla. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
T-REx: A large scale alignment of natural language with knowledge base triples
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frédérique Laforest, and Elena Simperl. 2018 · 2018
Cited alongside, same era.
Commonsense knowledge mining from pretrained models
Joe Davison, Joshua Feldman, and Alexander M Rush. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A Smith. 2019 · 2019
Closest in time.
Mask-Predict: Parallel decoding of conditional masked language models
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Closest in time.
Knowledge enhanced contextual word representations
Matthew E Peters, Mark Neumann, IV Logan, L Robert, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A Smith. 2019 · 2019
Closest in time.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
End-to-end neural entity linking
Nikolaos Kolitsas, Octavian-Eugen Ganea, and Thomas Hofmann. 2018 · 2018
Cited alongside, same era.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Cited alongside, same era.
Open domain question answering using early fusion of knowledge bases and text
Haitian Sun, Bhuwan Dhingra, Manzil Zaheer, Kathryn Mazaitis, Ruslan Salakhutdinov, and William Cohen. 2018 · 2018
Cited alongside, same era.
Investigating entity knowledge in bert with simple neural end-to-end entity linking
Samuel Broscheit. 2019 · 2019
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013a
Cited in the paper.
Exploiting similarities among languages for machine translation
Tomas Mikolov, Quoc V Le, and Ilya Sutskever. 2013b
Cited in the paper.
Improving pre-trained multilingual models with vocabulary expansion
Hai Wang, Dian Yu, Kai Sun, Janshu Chen, and Dong Yu. 2019a
Cited in the paper.
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Closest in time.
ERNIE: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2019
Closest in time.