Fetching the paper…
Reading the bibliography…
Pre-trained Language Models (PLMs) are trained on vast unlabeled data, rich in world knowledge.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
“Cloze Procedure”: A new tool for measuring readability
Wilson L. Taylor. 1953 · 1953
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t
Anna Gladkova, Aleksandr Drozd, and Satoshi Matsuoka. 2016 · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature
Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain Marshall, Ani Nenkova, and Byron Wallace. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
Joe Davison, Joshua Feldman, and Alexander Rush. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Earlier work this paper cites.
Inducing relational knowledge from BERT
Zied Bouraoui, Jose Camacho-Collados, and Steven Schockaert. 2020 · 2020
Earlier work this paper cites.
Pretrained language model embryology: The birth of ALBERT
Cheng-Han Chiang, Sung-Feng Huang, and Hung-yi Lee. 2020 · 2020
Earlier work this paper cites.
Entities as experts: Sparse memory access with entity supervision
Thibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski. 2020 · 2020
Earlier work this paper cites.
REALM: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Earlier work this paper cites.
EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering
Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020 · 2020
Earlier work this paper cites.
X-FACTR: Multilingual factual knowledge retrieval from pretrained language models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020a · 2020
Earlier work this paper cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020 · 2020
Earlier work this paper cites.
Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly
Nora Kassner and Hinrich Schütze. 2020 · 2020
Earlier work this paper cites.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
Exploring BERT’s sensitivity to lexical cues using tests from semantic priming
Kanishka Misra, Allyson Ettinger, and Julia Rayz. 2020 · 2020
Earlier work this paper cites.
E-BERT: Efficient-yet-effective entity embeddings for BERT
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Earlier work this paper cites.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
BERTnesia: Investigating the capture and forgetting of knowledge in BERT
Jonas Wallat, Jaspreet Singh, and Avishek Anand. 2020 · 2020
Earlier work this paper cites.
Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model
Wenhan Xiong, Jingfei Du, William Yang Wang, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
Probing pre-trained language models for disease knowledge
Israa Alghanmi, Luis Espinosa Anke, and Steven Schockaert. 2021 · 2021
Earlier work this paper cites.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021 · 2021
Earlier work this paper cites.
Perhaps PTLMs should go to school – a task to assess open book and closed book QA
Manuel Ciosici, Joe Cecil, Dong-Ho Lee, Alex Hedges, Marjorie Freedman, and Ralph Weischedel. 2021 · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Earlier work this paper cites.
Static embeddings as efficient knowledge bases?
Philipp Dufter, Nora Kassner, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
Prompt tuning or fine-tuning - investigating relational knowledge in pre-trained language models
Leandra Fichtel, Jan-Christoph Kalo, and Wolf-Tilo Balke. 2021 · 2021
Cited alongside, same era.
Pre-trained models: Past, present and future
Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Yuan Yao, Ao Zhang, Liang Zhang, Wentao Han, Minlie Huang, Qin Jin, Yanyan Lan, Yang Liu, Zhiyuan Liu, Zhiwu Lu, Xipeng Qiu, Ruihua Song, Jie Tang, Ji-Rong Wen, Jinhui Yuan, Wayne Xin Zhao, and Jun Zhu. 2021 · 2021
Cited alongside, same era.
BERTese: Learning to speak to BERT
Adi Haviv, Jonathan Berant, and Amir Globerson. 2021 · 2021
Cited alongside, same era.
Factual consistency of multilingual pretrained language models
Constanza Fierro and Anders Søgaard. 2022 · 2022
Later among the works it cites.
TemporalWiki: A lifelong benchmark for training and evaluating ever-evolving language models
Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, and Minjoon Seo. 2022a · 2022
Later among the works it cites.
KAMEL: Knowledge analysis with multitoken entities in language models
Jan-Christoph Kalo and Leandra Fichtel. 2022 · 2022
Later among the works it cites.
Plug-and-play adaptation for continuously-updated QA
Kyungjae Lee, Wookje Han, Seung-won Hwang, Hwaran Lee, Joonsuk Park, and Sang-Woo Lee. 2022 · 2022
Later among the works it cites.
ElitePLM: An empirical study on general language ability evaluation of pretrained language models
Junyi Li, Tianyi Tang, Zheng Gong, Lixin Yang, Zhuohao Yu, Zhipeng Chen, Jingyuan Wang, Xin Zhao, and Ji-Rong Wen. 2022a · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models as knowledge bases: On entity representations, storage capacity, and paraphrased queries
Benjamin Heinzerling and Kentaro Inui. 2021 · 2021
Cited alongside, same era.
Understanding by understanding not: Modeling negation in language models
Arian Hosseini, Siva Reddy, Dzmitry Bahdanau, R Devon Hjelm, Alessandro Sordoni, and Aaron Courville. 2021 · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2021 · 2021
Cited alongside, same era.
AMMUS : A survey of transformer-based pretrained models in natural language processing
Katikapalli Subramanyam Kalyan, Ajit Rajasekharan, and Sivanesan Sangeetha. 2021 · 2021
Cited alongside, same era.
Multilingual LAMA: Investigating knowledge in multilingual pretrained language models
Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021 · 2021
Cited alongside, same era.
Reordering examples helps during priming-based few-shot learning
Sawan Kumar and Partha Talukdar. 2021 · 2021
Cited alongside, same era.
Question and answer test-train overlap in open-domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel. 2021 · 2021
Cited alongside, same era.
How pre-trained language models capture factual knowledge? a causal-inspired analysis
Shaobo Li, Xiaoguang Li, Lifeng Shang, Zhenhua Dong, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang, and Qun Liu. 2022b · 2022
Later among the works it cites.
SPE: Symmetrical prompt enhancement for fact probing
Yiyuan Li, Tong Che, Yezhen Wang, Zhengbao Jiang, Caiming Xiong, and Snigdha Chaturvedi. 2022c · 2022
Later among the works it cites.
Coherence boosting: When your pretrained language model is not paying enough attention
Nikolay Malkin, Zhen Wang, and Nebojsa Jojic. 2022 · 2022
Later among the works it cites.
P-Adapters: Robustly extracting factual information from language models with diverse prompts
Benjamin Newman, Prafulla Kumar Choubey, and Nazneen Rajani. 2022 · 2022
Later among the works it cites.
Entity cloze by date: What LMs know about unseen entities
Yasumasa Onoe, Michael Zhang, Eunsol Choi, and Greg Durrett. 2022 · 2022
Later among the works it cites.
InforMask: Unsupervised informative masking for language model pretraining
Nafis Sadeq, Canwen Xu, and Julian McAuley. 2022 · 2022
Later among the works it cites.
You are my type! type embeddings for pre-trained language models
Mohammed Saeed and Paolo Papotti. 2022 · 2022
Later among the works it cites.
Knowledge base construction from pre-trained language models 2022
Sneha Singhania, Tuan-Phong Nguyen, and Simon Razniewski. 2022 · 2022
Later among the works it cites.
EntityCS: Improving zero-shot cross-lingual transfer with entity-centric code switching
Chenxi Whitehouse, Fenia Christopoulou, and Ignacio Iacobacci. 2022 · 2022
Later among the works it cites.
ZeroGen: Efficient zero-shot learning via dataset generation
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022 · 2022
Later among the works it cites.
PromptGen: Automatically generate prompts using generative models
Yue Zhang, Hongliang Fei, Dingcheng Li, and Ping Li. 2022 · 2022
Later among the works it cites.
On the explainability of natural language processing deep models
Julia El Zini and Mariette Awad. 2022 · 2022
Later among the works it cites.
The life cycle of knowledge in big language models: A survey
Boxi Cao, Hongyu Lin, Xianpei Han, and Le Sun. 2023 · 2023
Closest in time.
Salient span masking for temporal understanding
Jeremy R. Cole, Aditi Chaudhary, Bhuwan Dhingra, and Partha Talukdar. 2023 · 2023
Closest in time.
Methods for measuring, updating, and visualizing factual beliefs in language models
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2023 · 2023
Closest in time.
Detecting edit failures in large language models: An improved specificity benchmark
Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, and Fazl Barez. 2023 · 2023
Closest in time.
Evaluating the robustness of discrete prompts
Yoichi Ishibashi, Danushka Bollegala, Katsuhito Sudoh, and Satoshi Nakamura. 2023 · 2023
Closest in time.
Simple and effective multi-token completion from masked language models
Oren Kalinsky, Guy Kushilevitz, Alexander Libov, and Yoav Goldberg. 2023 · 2023
Closest in time.
Large language models struggle to learn long-tail knowledge
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023 · 2023
Closest in time.
DLAMA: A framework for curating culturally diverse facts for probing the knowledge of pretrained language models
Amr Keleg and Walid Magdy. 2023 · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Closest in time.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
Dynamic benchmarking of masked language models on temporal concept drift with multiple views
Katerina Margatina, Shuai Wang, Yogarshi Vyas, Neha Anna John, Yassine Benajiba, and Miguel Ballesteros. 2023 · 2023
Closest in time.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry. 2023 · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023 · 2023
Closest in time.
ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
Partha Pratim Ray. 2023 · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Towards alleviating the object bias in prompt tuning-based factual knowledge extraction
Yuhang Wang, Dongyuan Lu, Chao Kong, and Jitao Sang. 2023 · 2023
Closest in time.
Self-evolution learning for discriminative language model pretraining
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023 · 2023
Closest in time.
Selective-LAMA: Selective prediction for confidence-aware evaluation of language models
Hiyori Yoshikawa and Naoaki Okazaki. 2023 · 2028
Closest in time.
Nonparametric masked language modeling
Sewon Min, Weijia Shi, Mike Lewis, Xilun Chen, Wen-tau Yih, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2023 · 2097
Closest in time.