Fetching the paper…
Reading the bibliography…
Recently, there has been a surge of interest in the NLP community on the use of pretrained Language Models (LMs) as Knowledge Bases (KBs).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020a · 1901
Earlier work this paper cites.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019 · 1904
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019 · 1904
Earlier work this paper cites.
ERNIE: enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Yih, and Yejin Choi. 2020 · 1908
Earlier work this paper cites.
Designing and Interpreting Probes with Control Tasks
John Hewitt and Percy Liang. 2019 · 1909
Earlier work this paper cites.
Kg-bert: Bert for knowledge graph completion
Liang Yao, Chengsheng Mao, and Yuan Luo. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 1910
Earlier work this paper cites.
How Can We Know What Language Models Know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020b · 1911
Earlier work this paper cites.
oLMpics – On what Language Model Pre-training Captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020a · 1912
Earlier work this paper cites.
“cloze procedure”: A new tool for measuring readability
Wilson L Taylor. 1953 · 1953
Earlier work this paper cites.
The logic theory machine–a complex information processing system
A. Newell and H. Simon. 1956 · 1956
Earlier work this paper cites.
Programs with common sense
John McCarthy. 1959 · 1959
Earlier work this paper cites.
The influence curve and its role in robust estimation
Frank R. Hampel. 1974 · 1974
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Easy cases of probabilistic satisfiability
KimAllan Andersen and Daniele Pretolani. 2001 · 2001
Earlier work this paper cites.
Transformers as soft reasoners over language
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020 · 2002
Earlier work this paper cites.
REALM: Retrieval-Augmented Language Model Pre-Training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
A Survey on Knowledge Graphs: Representation, Acquisition and Applications
Shaoxiong Ji, Shirui Pan, Erik Cambria, Pekka Marttinen, and Philip S. Yu. 2021 · 2002
Earlier work this paper cites.
Expert systems in production planning and scheduling: A state-of-the-art survey
Kostas S. Metaxiotis, Dimitris Askounis, and John E. Psarras. 2002 · 2002
Earlier work this paper cites.
A Primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2002
Earlier work this paper cites.
Entities as experts: Sparse memory access with entity supervision
Thibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski. 2020 · 2004
Earlier work this paper cites.
Wt5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
A smorgasbord of features for statistical machine translation
Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alex Fraser, Shankar Kumar, Libin Shen, David Smith, Katherine Eng, Viren Jain, Zhen Jin, and Dragomir Radev. 2004 · 2004
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021 · 2005
Earlier work this paper cites.
Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit Knowledge
Alon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg, and Jonathan Berant. 2020b · 2006
Earlier work this paper cites.
Knowledge-aware language model pretraining
Corby Rosset, Chenyan Xiong, Minh Hieu Phan, Xia Song, Paul Bennett, and Saurabh Tiwary. 2020 · 2007
Earlier work this paper cites.
Facts as experts: Adaptable and interpretable neural memory over symbolic knowledge
Pat Verga, Haitian Sun, Livio Baldini Soares, and William W. Cohen. 2020 · 2007
Earlier work this paper cites.
Benjamin Heinzerling and Kentaro Inui. 2021 · 2008
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020 · 2009
Earlier work this paper cites.
Measuring systematic generalization in neural proof generation with transformers
Nicolas Gontier, Koustuv Sinha, Siva Reddy, and Christopher Joseph Pal. 2020 · 2009
Earlier work this paper cites.
Generative language modeling for automated theorem proving
Stanislas Polu and Ilya Sutskever. 2020 · 2009
Earlier work this paper cites.
A Survey of the State of Explainable AI for Natural Language Processing
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020 · 2010
Earlier work this paper cites.
Prover: Proof generation for interpretable reasoning over rules
Swarnadeep Saha, Sayan Ghosh, Shashank Srivastava, and Mohit Bansal. 2020 · 2010
Earlier work this paper cites.
Language Models are Open Knowledge Graphs
Chenguang Wang, Xiao Liu, and Dawn Song. 2020 · 2010
Earlier work this paper cites.
On the practical ability of recurrent neural networks to recognize hierarchical languages
S. Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2011
Earlier work this paper cites.
Measuring and repairing inconsistency in probabilistic knowledge bases
David Picado-Muiño. 2011 · 2011
Earlier work this paper cites.
Transition-based dependency parsing with rich non-local features
Yue Zhang and Joakim Nivre. 2011 · 2011
Earlier work this paper cites.
Extracting Training Data from Large Language Models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2012
Earlier work this paper cites.
Modifying Memories in Transformer Models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. 2020 · 2012
Earlier work this paper cites.
What Is a Paraphrase?
Rahul Bhagat and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Çaglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Wikidata: A free collaborative knowledgebase
Denny Vrandečić and Markus Krötzsch. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Jason Weston, Sumit Chopra, and Antoine Bordes. 2015 · 2015
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and T. Jaakkola. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
A neural knowledge language model
Sungjin Ahn, Heeyoul Choi, Tanel Pärnamaa, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
A survey on dialogue systems
Hongshen Chen, Xiaorui Liu, Dawei Yin, and Jiliang Tang. 2017 · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, P. Abbeel, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
Hypernetworks
David Ha, Andrew Dai, and Quoc V. Le. 2017 · 2017
Cited alongside, same era.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, undefinedukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
Combining pre-trained language models and structured knowledge
Pedro Colon-Hernandez, Catherine Havasi, Jason Alonso, Matthew Huggins, and Cynthia Breazeal. 2021 · 2021
Later among the works it cites.
On commonsense cues in BERT for solving commonsense tasks
Leyang Cui, Sijie Cheng, Yu Wu, and Yue Zhang. 2021 · 2021
Later among the works it cites.
Analyzing commonsense emergence in few-shot knowledge models
Jeff Da, Ronan Le Bras, Ximing Lu, Yejin Choi, and Antoine Bosselut. 2021 · 2021
Later among the works it cites.
Editing Factual Knowledge in Language Models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explainable artificial intelligence: A survey
Filip Karlo Došilović, Mario Brčić, and Nikica Hlupić. 2018 · 2018
Cited alongside, same era.
The mythos of model interpretability
Zachary Chase Lipton. 2018 · 2018
Cited alongside, same era.
Dissecting contextual word embeddings: Architecture and representation
Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen tau Yih. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Cited alongside, same era.
ParaNMT-50M: Pushing the limits of paraphrastic sentence embeddings with millions of machine translations
John Wieting and Kevin Gimpel. 2018 · 2018
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Mention memory: incorporating textual knowledge into transformers through entity mention attention
Michiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Fei Sha, and William Cohen. 2021 · 2021
Later among the works it cites.
Time-Aware Language Models as Temporal Knowledge Bases
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2021 · 2021
Later among the works it cites.
Measuring and Improving Consistency in Pretrained Language Models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Later among the works it cites.
Empowering language understanding with counterfactual reasoning
Fuli Feng, Jizhi Zhang, Xiangnan He, Hanwang Zhang, and Tat-Seng Chua. 2021 · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Later among the works it cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Later among the works it cites.
Ptr: Prompt tuning with rules for text classification
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Later among the works it cites.
Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2021 · 2021
Later among the works it cites.
Reasoning with transformer-based models: Deep learning, but shallow reasoning
Chadi Helwe, Chloé Clavel, and Fabian M. Suchanek. 2021 · 2021
Later among the works it cites.
BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of Belief
Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark. 2021 · 2021
Later among the works it cites.
Mind the Gap: Assessing Temporal Generalization in Neural Language Models
Angeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomas Kocisky, Sebastian Ruder, Dani Yogatama, Kris Cao, Susannah Young, and Phil Blunsom. 2021 · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
Causalbert: Injecting causal knowledge into pre-trained models with minimal supervision
Zhongyang Li, Xiao Ding, Kuo Liao, Ting Liu, and Bing Qin. 2021 · 2021
Later among the works it cites.
Towards understanding and mitigating social biases in language models
Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2021 · 2021
Later among the works it cites.
Cutting down on prompts and parameters: Simple few-shot learning with language models
Robert L Logan, Ivana Balavzevi’c, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel. 2021 · 2021
Later among the works it cites.
Multi-task retrieval for knowledge-intensive tasks
Jean Maillard, Vladimir Karpukhin, Fabio Petroni, Wen tau Yih, Barlas Oğuz, Veselin Stoyanov, and Gargi Ghosh. 2021 · 2021
Later among the works it cites.
Provable limitations of acquiring meaning from ungrounded form: What will future language models understand?
William Cooper Merrill, Yoav Goldberg, Roy Schwartz, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
On semantic cognition, inductive generalization, and language models
Kanishka Misra. 2021 · 2021
Later among the works it cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2021 · 2021
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021 · 2021
Later among the works it cites.
Prompting contrastive explanations for commonsense reasoning tasks
Bhargavi Paranjape, Julian Michael, Marjan Ghazvininejad, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021 · 2021
Later among the works it cites.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner. 2021 · 2021
Later among the works it cites.
Erica: Improving entity and relation understanding for pre-trained language models via contrastive learning
Yujia Qin, Yankai Lin, Ryuichi Takanobu, Zhiyuan Liu, Peng Li, Heng Ji, Minlie Huang, Maosong Sun, and Jie Zhou. 2021 · 2021
Later among the works it cites.
mluke: The power of entity representations in multilingual pretrained language models
Ryokan Ri, Ikuya Yamada, and Yoshimasa Tsuruoka. 2021 · 2021
Later among the works it cites.
Relational world knowledge representation in contextual language models: A review
Tara Safavi and Danai Koutra. 2021 · 2021
Later among the works it cites.
Interactively generating explanations for transformer language models
Patrick Schramowski, Felix Friedrich, Christopher Tauchmann, and Kristian Kersting. 2021 · 2021
Later among the works it cites.
Reasoning over virtual knowledge bases with open predicate relations
Haitian Sun, Pat Verga, Bhuwan Dhingra, Ruslan Salakhutdinov, and William W. Cohen. 2021 · 2021
Later among the works it cites.
Can language models be biomedical knowledge bases?
Mujeen Sung, Jinhyuk Lee, Sean Yi, Minji Jeon, Sungdong Kim, and Jaewoo Kang. 2021 · 2021
Later among the works it cites.
ProofWriter: Generating implications, proofs, and abductive statements over natural language
Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021 · 2021
Later among the works it cites.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Jonas Wallat, Jaspreet Singh, and Avishek Anand. 2021 · 2021
Later among the works it cites.
Knowledge enhanced pretrained language models: A compreshensive survey
Xiaokai Wei, Shen Wang, Dejiao Zhang, Parminder Bhatia, and Andrew Arnold. 2021 · 2021
Later among the works it cites.
Symbolic Knowledge Distillation: from General Language Models to Commonsense Models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Qa-gnn: Reasoning with language models and knowledge graphs for question answering
Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021 · 2021
Later among the works it cites.
Adaptive Semiparametric Language Models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Later among the works it cites.
Differentiable prompt makes pre-trained language models better few-shot learners
Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen. 2021 · 2021
Later among the works it cites.
Of non-linearity and commutativity in bert
Sumu Zhao, Damian Pascual, Gino Brunner, and Roger Wattenhofer. 2021 · 2021
Later among the works it cites.
Factual probing is [mask]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Later among the works it cites.
HILDIF: Interactive debugging of NLI models using influence functions
Hugo Zylberajch, Piyawat Lertvittayakumjorn, and Francesca Toni. 2021 · 2021
Later among the works it cites.
Cm3: A causal masked multimodal model of the internet
Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, and Luke Zettlemoyer. 2022 · 2022
Closest in time.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, P. Abbeel, Deepak Pathak, and Igor Mordatch. 2022 · 2022
Closest in time.
Rethinking explainability as a dialogue: A practitioner’s perspective
Himabindu Lakkaraju, Dylan Slack, Yuxin Chen, Chenhao Tan, and Sameer Singh. 2022 · 2022
Closest in time.
Relational memory augmented language models
Qi Liu, Dani Yogatama, and Phil Blunsom. 2022 · 2022
Closest in time.
Locating and editing factual knowledge in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Closest in time.
Formal mathematics statement curriculum learning
Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, and Ilya Sutskever. 2022 · 2022
Closest in time.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Closest in time.