Fetching the paper…
Reading the bibliography…
Pretrained Language Models (PLM) have established a new paradigm through learning informative contextualized representations on large-scale text corpus.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The unified medical language system (umls): integrating biomedical terminology
O. Bodenreider · 2004
Earlier work this paper cites.
Freebase: A collaboratively created graph database for structuring human knowledge
K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor · 2008
Earlier work this paper cites.
SentiWordNet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining
S. Baccianella, A. Esuli, and F. Sebastiani · 2010
Earlier work this paper cites.
Toward an architecture for never-ending language learning
A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. H. Jr., and T. Mitchell · 2010
Earlier work this paper cites.
Tagme: On-the-fly annotation of short text fragments (by wikipedia entities)
P. Ferragina and U. Scaiella · 2010
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
A. Bordes, N. Usunier, A. Garcia-Duran, J. Weston, and O. Yakhnenko · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Fast and accurate shift-reduce constituent parsing
M. Zhu, Y. Zhang, W. Chen, M. Zhang, and J. Zhu · 2013
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
D. Chen and C. Manning · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
SemEval-2014 task 4: Aspect based sentiment analysis
M. Pontiki, D. Galanis, J. Pavlopoulos, H. Papageorgiou, I. Androutsopoulos, and S. Manandhar · 2014
Earlier work this paper cites.
Wikidata: A free collaborative knowledgebase
D. Vrandečić and M. Krötzsch · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Design challenges for entity linking
X. Ling, S. Singh, and D. S. Weld · 2015
Earlier work this paper cites.
Inferring networks of substitutable and complementary products
J. McAuley, R. Pandey, and J. Leskovec · 2015
Earlier work this paper cites.
A corpus and cloze evaluation for deeper understanding of commonsense stories
N. Mostafazadeh, N. Chambers, X. He, D. Parikh, D. Batra, L. Vanderwende, P. Kohli, and J. Allen · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
M. Dunn, L. Sagun, M. Higgins, V. Ugur Guney, V. Cirik, and K. Cho · 2017
Earlier work this paper cites.
An end-to-end model for question answering over knowledge base with cross-attention combining global knowledge
Y. Hao, Y. Zhang, K. Liu, S. He, Z. Liu, H. Wu, and J. Zhao · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Conceptnet 5.5: An open multilingual graph of general knowledge
R. Speer, J. Chin, and C. Havasi · 2017
Earlier work this paper cites.
Know-evolve: Deep temporal reasoning for dynamic knowledge graphs
R. Trivedi, H. Dai, Y. Wang, and L. Song · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Position-aware attention and supervised data improve slot filling
Y. Zhang, V. Zhong, D. Chen, G. Angeli, and C. D. Manning · 2017
Earlier work this paper cites.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Y. Cheng, D. Wang, P. Zhou, and T. Zhang · 2018
Earlier work this paper cites.
Ultra-fine entity typing
E. Choi, O. Levy, Y. Choi, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Convolutional 2D knowledge graph embeddings
T. Dettmers, P. Minervini, P. Stenetorp, and S. Riedel · 2018
Earlier work this paper cites.
FewRel: A large-scale supervised few-shot relation classification dataset with state-of-the-art evaluation
X. Han, H. Zhu, P. Yu, Z. Wang, Y. Yao, Z. Liu, and M. Sun · 2018
Earlier work this paper cites.
Automated phrase mining from massive text corpora
J. Shang, J. Liu, M. Jiang, X. Ren, C. R. Voss, and J. Han · 2018
Earlier work this paper cites.
Graph attention networks
P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2018
Earlier work this paper cites.
COMET: Commonsense transformers for automatic knowledge graph construction
A. Bosselut, H. Rashkin, M. Sap, C. Malaviya, A. Celikyilmaz, and Y. Choi · 2019
Earlier work this paper cites.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Latent Relation Language Models
H. Hayashi, Z. Hu, C. Xiong, and G. Neubig · 2019
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
L. Huang, R. Le Bras, C. Bhagavatula, and Y. Choi · 2019
Earlier work this paper cites.
KagNet: Knowledge-aware graph networks for commonsense reasoning
B. Y. Lin, X. Chen, J. Chen, and X. Ren · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
N. F. Liu, M. Gardner, Y. Belinkov, M. E. Peters, and N. A. Smith · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Barack’s wife hillary: Using knowledge graphs for fact-aware language modeling
R. Logan, N. F. Liu, M. E. Peters, M. Gardner, and S. Singh · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
T. McCoy, E. Pavlick, and T. Linzen · 2019
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2020
Later among the works it cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela · 2020
Later among the works it cites.
Lexically-constrained Text Generation through Commonsense Knowledge Extraction and Injection
Y. Li, P. Goel, V. Kuppur Rajendra, H. Simrat Singh, J. Francis, K. Ma, E. Nyberg, and A. Oltramari · 2020
Later among the works it cites.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
B. Y. Lin, W. Zhou, M. Shen, P. Zhou, C. Bhagavatula, Y. Choi, and X. Ren · 2020
Later among the works it cites.
K-BERT: enabling language representation with knowledge graph
W. Liu, P. Zhou, Z. Zhao, Z. Wang, Q. Ju, H. Deng, and P. Wang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probing neural network comprehension of natural language arguments
T. Niven and H.-Y. Kao · 2019
Cited alongside, same era.
Enriching BERT with Knowledge Graph Embedding for Document Classification
M. Ostendorff, P. Bourgonje, M. Berger, J. Moreno-Schneider, and G. Rehm · 2019
Cited alongside, same era.
Knowledge enhanced contextual word representations
M. E. Peters, M. Neumann, R. Logan, R. Schwartz, V. Joshi, S. Singh, and N. A. Smith · 2019
Cited alongside, same era.
Language models as knowledge bases?
F. Petroni, T. Rocktäschel, S. Riedel, P. Lewis, A. Bakhtin, Y. Wu, and A. Miller · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Cited alongside, same era.
ATOMIC: an atlas of machine commonsense for if-then reasoning
M. Sap, R. L. Bras, E. Allaway, C. Bhagavatula, N. Lourie, H. Rashkin, B. Roof, N. A. Smith, and Y. Choi · 2019
Cited alongside, same era.
Later among the works it cites.
KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense Reasoning
Y. Liu, Y. Wan, L. He, H. Peng, and P. S. Yu · 2020
Later among the works it cites.
Graph-based reasoning over heterogeneous external knowledge for commonsense question answering
S. Lv, D. Guo, J. Xu, D. Tang, N. Duan, M. Gong, L. Shou, D. Jiang, G. Cao, and S. Hu · 2020
Later among the works it cites.
The effect of natural distribution shift on question answering models
J. Miller, K. Krauth, B. Recht, and L. Schmidt · 2020
Later among the works it cites.
How context affects language models’ factual predictions
F. Petroni, P. Lewis, A. Piktus, T. Rocktäschel, Y. Wu, A. H. Miller, and S. Riedel · 2020
Later among the works it cites.
E-BERT: Efficient-yet-effective entity embeddings for BERT
N. Poerner, U. Waltinger, and H. Schütze · 2020
Later among the works it cites.
Y. Qin, Y. Lin, R. Takanobu, Z. Liu, P. Li, H. Ji, M. Huang, M. Sun, and J. Zhou · 2020
Later among the works it cites.
Pre-trained models for natural language processing: A survey
X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
A. Roberts, C. Raffel, and N. Shazeer · 2020
Later among the works it cites.
Knowledge-Aware Language Model Pretraining
C. Rosset, C. Xiong, M. Phan, X. Song, P. Bennett, and S. Tiwary · 2020
Later among the works it cites.
Q-BERT: hessian based ultra low precision quantization of BERT
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer · 2020
Later among the works it cites.
Exploiting structured knowledge in text via graph-guided representation learning
T. Shen, Y. Mao, P. He, G. Long, A. Trischler, and W. Chen · 2020
Later among the works it cites.
CokeBERT: Contextual Knowledge Selection and Embedding towards Enhanced Pre-Trained Language Models
Y. Su, X. Han, Z. Zhang, P. Li, Z. Liu, Y. Lin, J. Zhou, and M. Sun · 2020
Later among the works it cites.
Colake: Contextualized language and knowledge embedding
T. Sun, Y. Shao, X. Qiu, Q. Guo, Y. Hu, X. Huang, and Z. Zhang · 2020
Later among the works it cites.
ERNIE 2.0: A continual pre-training framework for language understanding
Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu, and H. Wang · 2020
Later among the works it cites.
oLMpics-on what language model pre-training captures
A. Talmor, Y. Elazar, Y. Goldberg, and J. Berant · 2020
Later among the works it cites.
SKEP: Sentiment knowledge enhanced pre-training for sentiment analysis
H. Tian, C. Gao, X. Xiao, H. Liu, B. He, H. Wu, H. Wang, and F. Wu · 2020
Later among the works it cites.
Facts as Experts: Adaptable and Interpretable Neural Memory over Symbolic Knowledge
P. Verga, H. Sun, L. Baldini Soares, and W. W. Cohen · 2020
Later among the works it cites.
Probing pretrained language models for lexical semantics
I. Vulić, E. M. Ponti, R. Litschko, G. Glavaš, and A. Korhonen · 2020
Later among the works it cites.
Language Models are Open Knowledge Graphs
C. Wang, X. Liu, and D. Song · 2020
Later among the works it cites.
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters
R. Wang, D. Tang, N. Duan, Z. Wei, X. Huang, J. ji, G. Cao, D. Jiang, and M. Zhou · 2020
Later among the works it cites.
Contrastive Learning for Sequential Recommendation
X. Xie, F. Sun, Z. Liu, S. Wu, J. Gao, B. Ding, and B. Cui · 2020
Later among the works it cites.
Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model
W. Xiong, J. Du, W. Y. Wang, and V. Stoyanov · 2020
Later among the works it cites.
Luke: Deep contextualized entity representations with entity-aware self-attention
I. Yamada, A. Asai, H. Shindo, H. Takeda, and Y. Matsumoto · 2020
Later among the works it cites.
CoCoLM: COmplex COmmonsense Enhanced Language Model
C. Yu, H. Zhang, Y. Song, and W. Ng · 2020
Later among the works it cites.
JAKET: Joint Pre-training of Knowledge Graph and Language Understanding
D. Yu, C. Zhu, Y. Yang, and M. Zeng · 2020
Later among the works it cites.
E-BERT: A Phrase and Product Knowledge Enhanced Language Model for E-commerce
D. Zhang, Z. Yuan, Y. Liu, Z. Fu, F. Zhuang, P. Wang, H. Chen, and H. Xiong · 2020
Later among the works it cites.
LIMIT-BERT : Linguistics informed multi-task BERT
J. Zhou, Z. Zhang, H. Zhao, and S. Zhang · 2020
Later among the works it cites.
Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees
J. Bai, Y. Wang, Y. Chen, Y. Yang, J. Bai, J. Yu, and Y. Tong · 2021
Closest in time.
Combining pre-trained language models and structured knowledge
P. Colon-Hernandez, C. Havasi, J. Alonso, M. Huggins, and C. Breazeal · 2021
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2021
Closest in time.
Relational world knowledge representation in contextual language models: A review
T. Safavi and D. Koutra · 2021
Closest in time.