Fetching the paper…
Reading the bibliography…
Pretrained language models (LMs) do not capture factual knowledge very well.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
KG-BERT: BERT for knowledge graph completion
Liang Yao, Chengsheng Mao, and Yuan Luo. 2019 · 1909
Earlier work this paper cites.
The Fourier transform and its applications , volume 31999
Ronald Newbold Bracewell and Ronald N Bracewell. 1986 · 1986
Earlier work this paper cites.
Approximation by superpositions of a sigmoidal function
George Cybenko. 1992 · 1992
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger. 2004 · 2004
Earlier work this paper cites.
On bayesian bounds
Arindam Banerjee. 2006 · 2006
Earlier work this paper cites.
Understanding graph neural networks from graph signal denoising perspectives
Guoji Fu, Yifan Hou, Jian Zhang, Kaili Ma, Barakeel Fanseu Kamhoua, and James Cheng. 2020 · 2006
Earlier work this paper cites.
TAGME: on-the-fly annotation of short text fragments (by wikipedia entities)
Paolo Ferragina and Ugo Scaiella. 2010 · 2010
Earlier work this paper cites.
Interpreting graph neural networks for NLP with differentiable edge masking
Michael Sejr Schlichtkrull, Nicola De Cao, and Ivan Titov. 2020 · 2010
Earlier work this paper cites.
Link prediction via matrix factorization
Aditya Krishna Menon and Charles Elkan. 2011 · 2011
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Oksana Yakhnenko. 2013 · 2013
Earlier work this paper cites.
Spectral networks and locally connected networks on graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann LeCun. 2014 · 2014
Earlier work this paper cites.
Discrete signal processing on graphs: Frequency analysis
Aliaksei Sandryhaila and José M. F. Moura. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Design challenges for entity linking
Xiao Ling, Sameer Singh, and Daniel S. Weld. 2015 · 2015
Earlier work this paper cites.
Convolutional neural networks on graphs with fast localized spectral filtering
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. 2016 · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil C. Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. 2016 · 2016
Earlier work this paper cites.
Handling class imbalance in link prediction using learning to rank techniques
Bopeng Li, Sougata Chaudhuri, and Ambuj Tewari. 2016 · 2016
Earlier work this paper cites.
“why should I trust you?”: Explaining the predictions of any classifier
Marco Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling. 2017 · 2017
Cited alongside, same era.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, R. Devon Hjelm, and Aaron C. Courville. 2018 · 2018
Cited alongside, same era.
Ultra-fine entity typing
Eunsol Choi, Omer Levy, Yejin Choi, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
T-REx: A large scale alignment of natural language with knowledge base triples
Hady Elsahar, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon Hare, Frederique Laforest, and Elena Simperl. 2018 · 2018
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Later among the works it cites.
Do NLP models know numbers? probing numeracy in embeddings
Eric Wallace, Yizhong Wang, Sujian Li, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
ERNIE: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2019
Later among the works it cites.
K-BERT: enabling language representation with knowledge graph
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Qi Ju, Haotang Deng, and Ping Wang. 2020 · 2020
Later among the works it cites.
Information-theoretic probing for linguistic structure
Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modeling relational data with graph convolutional networks
Michael Sejr Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. 2018 · 2018
Cited alongside, same era.
Attention-based graph neural network for semi-supervised learning
Kiran Koshy Thekumparampil, Chong Wang, Sewoong Oh, and Li-Jia Li. 2018 · 2018
Cited alongside, same era.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
On the equivalence between graph isomorphism testing and function approximation with gnns
Zhengdao Chen, Soledad Villar, Lei Chen, and Joan Bruna. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Orthogonal relation transforms with graph context modeling for knowledge graph embedding
Yun Tang, Jing Huang, Guangtao Wang, Xiaodong He, and Bowen Zhou. 2020 · 2020
Later among the works it cites.
Document modeling with graph attention networks for multi-grained machine reading comprehension
Bo Zheng, Haoyang Wen, Yaobo Liang, Nan Duan, Wanxiang Che, Daxin Jiang, Ming Zhou, and Ting Liu. 2020 · 2020
Later among the works it cites.
Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training
Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou. 2021 · 2021
Later among the works it cites.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021 · 2021
Later among the works it cites.
Time-aware language models as temporal knowledge bases
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2021 · 2021
Later among the works it cites.
Bird’s eye: Probing for linguistic graph structures with a simple information-theoretic approach
Yifan Hou and Mrinmaya Sachan. 2021 · 2021
Later among the works it cites.
How can we know when language models know? on the calibration of language models for question answering
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Prakhar Kaushik, Alex Gain, Adam Kortylewski, and Alan L. Yuille. 2021 · 2021
Later among the works it cites.
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou. 2021a · 2021
Later among the works it cites.
Factual probing is [MASK]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Temporal reasoning on implicit events from distant supervision
Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, and Dan Roth. 2021 · 2021
Later among the works it cites.
Probing via prompting
Jiaoda Li, Ryan Cotterell, and Mrinmaya Sachan. 2022 · 2022
Closest in time.