Fetching the paper…
Reading the bibliography…
Probing complex language models has recently revealed several insights into linguistic and semantic patterns found in the learned representations.
Assessing bert’s syntactic abilities
Goldberg, Yoav. 2019 · 1901
Earlier work this paper cites.
Nogueira, Rodrigo and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Sun, Yu, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019 · 1904
Earlier work this paper cites.
K-bert: Enabling language representation with knowledge graph
Liu, Weijie, Peng Zhou, Zhe Zhao, Zhiruo Wang, Q. Ju, Haotang Deng, and P. Wang. 2020 · 1909
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, Thomas, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, Michael and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Guu, Kelvin, Kenton Lee, Z. Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
A primer in bertology: What we know about how BERT works
Rogers, Anna, Olga Kovaleva, and Anna Rumshisky. 2020 · 2002
Earlier work this paper cites.
K-adapter: Infusing knowledge into pre-trained models with adapters
Wang, Ruize, Duyu Tang, Nan Duan, Zhongyu Wei, X. Huang, Jianshu Ji, Cuihong Cao, Daxin Jiang, and M. Zhou. 2020 · 2002
Earlier work this paper cites.
Deebert: Dynamic early exiting for accelerating bert inference
Xin, Ji, Raphael Tang, J. Lee, Y. Yu, and Jimmy Lin. 2020 · 2004
Earlier work this paper cites.
On the stability of fine-tuning BERT: misconceptions, explanations, and strong baselines
Mosbach, Marius, Maksym Andriushchenko, and Dietrich Klakow. 2020 · 2006
Earlier work this paper cites.
Kilt: a benchmark for knowledge intensive language tasks
Petroni, Fabio, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al. 2020 · 2009
Earlier work this paper cites.
Pmi-masking: Principled masking of correlated spans
Levine, Yoav, Barak Lenz, Opher Lieber, Omri Abend, Kevin Leyton-Brown, Moshe Tennenholtz, and Y. Shoham. 2020 · 2010
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, Mor, R. Schuster, Jonathan Berant, and Omer Levy. 2020 · 2012
Earlier work this paper cites.
Representing general relational knowledge in conceptnet 5
Speer, R. and Catherine Havasi. 2012 · 2012
Earlier work this paper cites.
50,000 lessons on how to read: a relation extraction corpus
Orr, Dave. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, Jeffrey, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Yukun, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
MS MARCO: A human generated machine reading comprehension dataset
Nguyen, Tri, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016 · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Rajpurkar, Pranav, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Joint learning of the embedding of words and entities for named entity disambiguation
Yamada, Ikuya, Hiroyuki Shindo, Hideaki Takeda, and Yoshiyasu Takefuji. 2016 · 2016
Earlier work this paper cites.
Exploring web archives through temporal anchor texts
Holzmann, Helge, Wolfgang Nejdl, and Avishek Anand. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Sundararajan, M., Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Evaluating compositionality in sentence embeddings
Dasgupta, Ishita, Demi Guo, Andreas Stuhlmüller, Samuel Gershman, and Noah D. Goodman. 2018 · 2018
Cited alongside, same era.
T-rex: A large scale alignment of natural language with knowledge base triples
ElSahar, Hady, Pavlos Vougiouklis, Arslen Remaci, Christophe Gravier, Jonathon S. Hare, Frédérique Laforest, and Elena Simperl. 2018 · 2018
Cited alongside, same era.
Assessing composition in sentence vector representations
Ettinger, Allyson, Ahmed Elgohary, Colin Phillips, and Philip Resnik. 2018 · 2018
Cited alongside, same era.
Language models as knowledge bases?
Petroni, Fabio, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019 · 2019
Later among the works it cites.
EXS: explainable search using local model agnostic interpretability
Singh, Jaspreet and Avishek Anand. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Tenney, Ian, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Tenney, Ian, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Inducing relational knowledge from BERT
Bouraoui, Zied, José Camacho-Collados, and Steven Schockaert. 2020 · 2020
Later among the works it cites.
BERT-MK: Integrating graph contextualized knowledge into pre-trained language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kiesel, Johannes, Arefeh Bahrami, Benno Stein, Avishek Anand, and Matthias Hagen. 2018 · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Peters, Matthew E., Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018 · 2018
Cited alongside, same era.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, Pranav, Robin Jia, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Posthoc interpretability of learning to rank models using secondary training data
Singh, Jaspreet and Avishek Anand. 2018 · 2018
Cited alongside, same era.
Sena-cnn: Overcoming catastrophic forgetting in convolutional neural networks by selective network augmentation
Zacarias, Abel S. and Luís A. Alexandre. 2018 · 2018
Cited alongside, same era.
How does BERT answer questions?: A layer-wise analysis of transformer representations
van Aken, Betty, Benjamin Winter, Alexander Löser, and Felix A. Gers. 2019 · 2019
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Belinkov, Yonatan and James R. Glass. 2019 · 2019
Cited alongside, same era.
He, Bin, Di Zhou, Jinghui Xiao, Xin Jiang, Qun Liu, Nicholas Jing Yuan, and Tong Xu. 2020 · 2020
Later among the works it cites.
What do compressed deep neural networks forget
Hooker, Sara, Aaron C. Courville, Gregory Clark, Yann Dauphin, and Andrea Frome. 2020 · 2020
Later among the works it cites.
Are pretrained language models symbolic reasoners over knowledge?
Kassner, Nora, Benno Krojer, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
E-BERT: Efficient-yet-effective entity embeddings for BERT
Poerner, Nina, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Roberts, Adam, Colin Raffel, and Noam Shazeer. 2020 · 2020
Later among the works it cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Shin, Taylor, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Singh, Jaspreet, Jonas Wallat, and Avishek Anand. 2020 · 2020
Later among the works it cites.
Knowledge neurons in pretrained transformers
Dai, Damai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2021 · 2021
Closest in time.
Measuring and improving consistency in pretrained language models
Elazar, Yanai, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, E. Hovy, H. Schutze, and Yoav Goldberg. 2021 · 2021
Closest in time.
Multilingual LAMA: Investigating knowledge in multilingual pretrained language models
Kassner, Nora, Philipp Dufter, and Hinrich Schütze. 2021 · 2021
Closest in time.
Enriching a model’s notion of belief using a persistent memory
Kassner, Nora, Oyvind Tafjord, H. Schutze, and P. Clark. 2021 · 2021
Closest in time.
Paq: 65 million probably-asked questions and what you can do with them
Lewis, Patrick, Yuxiang Wu, L. Liu, Pasquale Minervini, Heinrich Kuttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Closest in time.
Linear transformers are secretly fast weight memory systems
Schlag, Imanol, Kazuki Irie, and J. Schmidhuber. 2021 · 2021
Closest in time.
Not all memories are created equal: Learning to forget by expiring
Sukhbaatar, Sainbayar, Da Ju, Spencer Poff, Stephen Roller, Arthur D. Szlam, J. Weston, and Angela Fan. 2021 · 2021
Closest in time.
BERxiT: Early exiting for BERT with better fine-tuning and extension to regression
Xin, Ji, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021 · 2021
Closest in time.
Explain and predict, and then predict again
Zhang, Zijian, Koustav Rudra, and Avishek Anand. 2021 · 2021
Closest in time.
Factual probing is [mask]: Learning vs. learning to recall
Zhong, Zexuan, Dan Friedman, and Danqi Chen. 2021 · 2021
Closest in time.