Fetching the paper…
Reading the bibliography…
This paper surveys and organizes research works in a new paradigm in natural language processing, which we dub "prompt-based learning".
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019b · 1904
Earlier work this paper cites.
UNIFIEDQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, M. Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Zero-shot text classification with generative language models
Raul Puri and Bryan Catanzaro. 2019 · 1912
Earlier work this paper cites.
INTERFACILE: Linguistic coverage and query reformulation
Yvette Mathieu and Paul Sabatier. 1986 · 1986
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Measurement, regression, and calibration
Leon Jay Gleser. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Stacked generalizations: When does it work?
Kai Ming Ting and Ian H. Witten. 1997 · 1997
Earlier work this paper cites.
A bit of progress in language modeling
Joshua T Goodman. 2001 · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
J. Lafferty, A. McCallum, and Fernando Pereira. 2001 · 2001
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. 2020 · 2002
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
Gene selection for cancer classification using support vector machines
Isabelle Guyon, Jason Weston, Stephen Barnhill, and Vladimir Vapnik. 2002 · 2002
Earlier work this paper cites.
Thumbs up? sentiment classification using machine learning techniques
Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002 · 2002
Earlier work this paper cites.
Ensembling neural networks: many could be better than all
Zhi-Hua Zhou, Jianxin Wu, and Wei Tang. 2002 · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Abstractive summarization with combination of pre-trained sequence-to-sequence and saliency models
Itsumi Saito, Kyosuke Nishida, Kosuke Nishida, and Junji Tomita. 2020 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Enhancing factual consistency of abstractive summarization
Chenguang Zhu, William Hinthorn, Ruochen Xu, Qingkai Zeng, Michael Zeng, Xuedong Huang, and Meng Jiang. 2020 · 2003
Earlier work this paper cites.
Web search intent induction via automatic query reformulation
Hal Daumé III and Eric Brill. 2004 · 2004
Earlier work this paper cites.
A smorgasbord of features for statistical machine translation
Franz Josef Och, Daniel Gildea, Sanjeev Khudanpur, Anoop Sarkar, Kenji Yamada, Alex Fraser, Shankar Kumar, Libin Shen, David Smith, Katherine Eng, Viren Jain, Zhen Jin, and Dragomir Radev. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
How context affects language models’ factual predictions
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020 · 2005
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020b · 2006
Earlier work this paper cites.
An example of statistical investigation of the text eugene onegin concerning the connection of samples in chains
Andreĭ Andreevich Markov. 2006 · 2006
Earlier work this paper cites.
Supervised machine learning: A review of classification techniques
Sotiris B Kotsiantis, I Zaharakis, P Pintelas, et al. 2007 · 2007
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn. 2009 · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
A survey of knowledge-enhanced text generation
Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. 2020 · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, J. Weston, L. Bottou, Michael Karlen, K. Kavukcuoglu, and P. Kuksa. 2011 · 2011
Earlier work this paper cites.
Generalized minimum bayes risk system combination
Kevin Duh, Katsuhito Sudoh, Xianchao Wu, Hajime Tsukada, and Masaaki Nagata. 2011 · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector J. Levesque. 2011 · 2011
Earlier work this paper cites.
Transition-based dependency parsing with rich non-local features
Yue Zhang and Joakim Nivre. 2011 · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li. 2020 · 2012
Earlier work this paper cites.
Ctrlsum: Towards generic controllable text summarization
Junxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Fatema Rajani, and Caiming Xiong. 2020a · 2012
Earlier work this paper cites.
How can we know when language models know?
Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. 2020b · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Xuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2012
Earlier work this paper cites.
CPM: A large-scale generative chinese pre-trained language model
Zhengyan Zhang, Xu Han, Hao Zhou, Pei Ke, Yuxian Gu, Deming Ye, Yujia Qin, YuSheng Su, Haozhe Ji, Jian Guan, Fanchao Qi, Xiaozhi Wang, Yanan Zheng, Guoyang Zeng, Huanqi Cao, Shengqi Chen, Daixuan Li, Zhenbo Sun, Zhiyuan Liu, Minlie Huang, Wentao Han, Jie Tang, Juanzi Li, Xiaoyan Zhu, and Maosong Sun. 2020c · 2012
Earlier work this paper cites.
Computational approaches to sentence completion
Geoffrey Zweig, John C. Platt, Christopher Meek, Christopher J.C. Burges, Ainur Yessenalina, and Qiang Liu. 2012 · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent. 2013 · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
Identifying web search query reformulation using concept based matching
Ahmed Hassan. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Semantic parsing via paraphrasing
Jonathan Berant and Percy Liang. 2014 · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, Çaglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, R. Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Salicon: Saliency in context
Ming Jiang, Shengsheng Huang, Juanyong Duan, and Qi Zhao. 2015 · 2015
Earlier work this paper cites.
Learning feed-forward one-shot learners
Luca Bertinetto, João F Henriques, Jack Valmadre, Philip Torr, and Andrea Vedaldi. 2016 · 2016
Earlier work this paper cites.
Controlling output length in neural encoder-decoders
Yuta Kikuchi, Graham Neubig, Ryohei Sasano, Hiroya Takamura, and Manabu Okumura. 2016 · 2016
Earlier work this paper cites.
Ask me anything: Dynamic memory networks for natural language processing
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
Neural machine translation with supervised attention
Lemao Liu, Masao Utiyama, Andrew Finch, and Eiichiro Sumita. 2016 · 2016
Earlier work this paper cites.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard Hovy. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Controlling politeness in neural machine translation via side constraints
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a · 2016
Earlier work this paper cites.
Seeing with humans: Gaze-assisted neural image captioning
Yusuke Sugano and Andreas Bulling. 2016 · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanislaw Jastrzebski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al. 2017 · 2017
Cited alongside, same era.
Ask the right questions: Active question reformulation with reinforcement learning
Christian Buck, Jannis Bulian, Massimiliano Ciaramita, Wojciech Gajewski, Andrea Gesmundo, Neil Houlsby, and Wei Wang. 2017 · 2017
Cited alongside, same era.
An empirical comparison of domain adaptation methods for neural machine translation
Chenhui Chu, Raj Dabre, and Sadao Kurohashi. 2017 · 2017
Cited alongside, same era.
Data augmentation for low-resource neural machine translation
Marzieh Fadaee, Arianna Bisazza, and Christof Monz. 2017 · 2017
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Later among the works it cites.
Pre-trained models for natural language processing: A survey
Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Automatically identifying words that can serve as labels for few-shot text classification
Timo Schick, Helmut Schmid, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Timo Schick and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017a · 2017
Cited alongside, same era.
Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
Chuang Gan, Yandong Li, Haoxiang Li, Chen Sun, and Boqing Gong. 2017 · 2017
Cited alongside, same era.
Learning to remember rare events
Łukasz Kaiser, Ofir Nachum, Aurko Roy, and Samy Bengio. 2017 · 2017
Cited alongside, same era.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Cited alongside, same era.
End-to-end neural coreference resolution
Kenton Lee, Luheng He, Mike Lewis, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Task-oriented query reformulation with reinforcement learning
Rodrigo Nogueira and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Learning to compose domain-specific transformations for data augmentation
Alexander J. Ratner, Henry R. Ehrenberg, Zeshan Hussain, Jared Dunnmon, and Christopher Ré. 2017 · 2017
Cited alongside, same era.
Timo Schick and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Improving natural language processing tasks with human gaze-guided neural attention
Ekta Sood, Simon Tannert, Philipp Mueller, and Andreas Bulling. 2020 · 2020
Later among the works it cites.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
ERNIE 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
A wrong answer or a wrong question? an intricate relationship between question reformulation and answer selection in conversational question answering
Svitlana Vakulenko, Shayne Longpre, Zhucheng Tu, and Raviteja Anantha. 2020 · 2020
Later among the works it cites.
Generalizing from a few examples: A survey on few-shot learning
Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. 2020 · 2020
Later among the works it cites.
CorefQA: Coreference resolution as query-based span prediction
Wei Wu, Fei Wang, Arianna Yuan, Fei Wu, and Jiwei Li. 2020 · 2020
Later among the works it cites.
Designing templates for eliciting commonsense knowledge from pretrained sequence-to-sequence models
Jheng-Hong Yang, Sheng-Chieh Lin, Rodrigo Nogueira, Ming-Feng Tsai, Chuan-Ju Wang, and Jimmy Lin. 2020 · 2020
Later among the works it cites.
Tabert: Pretraining for joint understanding of textual and tabular data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Later among the works it cites.
PEGASUS: pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020a · 2020
Later among the works it cites.
Human gaze assisted artificial intelligence: a review
Ruohan Zhang, Akanksha Saran, Bo Liu, Yifeng Zhu, Sihang Guo, Scott Niekum, Dana Ballard, and Mary Hayhoe. 2020b · 2020
Later among the works it cites.
Htlm: Hyper-text pre-training and prompting of language models
Armen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi, Hu Xu, Gargi Ghosh, and Luke Zettlemoyer. 2021 · 2021
Closest in time.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei. 2021 · 2021
Closest in time.
Pada: A prompt-based autoregressive approach for adaptation to unseen domains
Eyal Ben-David, Nadav Oved, and Roi Reichart. 2021 · 2021
Closest in time.
GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021 · 2021
Closest in time.
FLEX: unifying evaluation for few-shot NLP
Jonathan Bragg, Arman Cohan, Kyle Lo, and Iz Beltagy. 2021 · 2021
Closest in time.
Template-based named entity recognition using bart
Leyang Cui, Yu Wu, Jian Liu, Sen Yang, and Yue Zhang. 2021 · 2021
Closest in time.
A primer on pretrained multilingual language models
Sumanth Doddapaneni, Gowtham Ramesh, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh M Khapra. 2021 · 2021
Closest in time.
GSum: A general framework for guided neural abstractive summarization
Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, and Graham Neubig. 2021 · 2021
Closest in time.
All nlp tasks are generation tasks: A general pretraining framework
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2021 · 2021
Closest in time.
Spanner: Named entity re-/recognition as span prediction
Jinlan Fu, Xuanjing Huang, and Pengfei Liu. 2021 · 2021
Closest in time.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Closest in time.
Warp: Word-level adversarial reprogramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Closest in time.
Ptr: Prompt tuning with rules for text classification
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Closest in time.
BERTese: Learning to speak to BERT
Adi Haviv, Jonathan Berant, and Amir Globerson. 2021 · 2021
Closest in time.
Surface form competition: Why the highest probability answer isn’t always right
Ari Holtzman, Peter West, Vered Schwartz, Yejin Choi, and Luke Zettlemoyer. 2021 · 2021
Closest in time.
Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition
Daniel Jurafsky and James H Martin. 2021 · 2021
Closest in time.
Reordering examples helps during priming-based few-shot learning
Sawan Kumar and Partha Talukdar. 2021 · 2021
Closest in time.
How many data points is a prompt worth?
Teven Le Scao and Alexander Rush. 2021 · 2021
Closest in time.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Closest in time.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Closest in time.
RefSum: Refactoring neural summarization
Yixin Liu, Zi-Yi Dou, and Pengfei Liu. 2021c · 2021
Closest in time.
Simcls: A simple framework for contrastive learning of abstractive summarization
Yixin Liu and Pengfei Liu. 2021 · 2021
Closest in time.
Cutting down on prompts and parameters: Simple few-shot learning with language models
Robert L. Logan IV, Ivana Balažević, Eric Wallace, Fabio Petroni, Sameer Singh, and Sebastian Riedel. 2021 · 2021
Closest in time.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, A. Moore, S. Riedel, and Pontus Stenetorp. 2021 · 2021
Closest in time.
Natural instructions: Benchmarking generalization to new tasks from natural language instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2021 · 2021
Closest in time.
True few-shot learning with language models
Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021 · 2021
Closest in time.
Learning how to ask: Querying LMs with mixtures of soft prompts
Guanghui Qin and Jason Eisner. 2021 · 2021
Closest in time.
Prompt programming for large language models: Beyond the few-shot paradigm
Laria Reynolds and Kyle McDonell. 2021 · 2021
Closest in time.
A mathematical exploration of why language models help solve downstream tasks
Nikunj Saunshi, Sadhika Malladi, and Sanjeev Arora. 2021 · 2021
Closest in time.
How many data points is a prompt worth?
Teven Le Scao and Alexander M Rush. 2021 · 2021
Closest in time.
Generating datasets with pretrained language models
Timo Schick and Hinrich Schütze. 2021 · 2021
Closest in time.
Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp
Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021 · 2021
Closest in time.
Constrained language models yield few-shot semantic parsers
Richard Shin, C. H. Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, D. Klein, J. Eisner, and Benjamin Van Durme. 2021 · 2021
Closest in time.
ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, Weixin Liu, Zhihua Wu, Weibao Gong, Jianzhong Liang, Zhizhou Shang, Peng Sun, Wei Liu, Xuan Ouyang, Dianhai Yu, Hao Tian, Hua Wu, and Haifeng Wang. 2021 · 2021
Closest in time.
Improving and simplifying pattern exploiting training
Derek Tam, Rakesh R Menon, Mohit Bansal, Shashank Srivastava, and Colin Raffel. 2021 · 2021
Closest in time.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Closest in time.
KEPLER: A unified model for knowledge embedding and pre-trained language representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang. 2021 · 2021
Closest in time.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
Colin Wei, Sang Michael Xie, and Tengyu Ma. 2021 · 2021
Closest in time.
Ernie-gram: Pre-training with explicitly n-gram masked language modeling for natural language understanding
Dongling Xiao, Yu-Kun Li, Han Zhang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021 · 2021
Closest in time.
Pre-trained models: Past, present and future
Han Xu, Zhang Zhengyan, Ding Ning, Gu Yuxian, Liu Xiao, Huo Yuqi, Qiu Jiezhong, Zhang Liang, Han Wentao, Huang Minlie, et al. 2021 · 2021
Closest in time.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021b · 2021
Closest in time.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, Chen Li, Ziyan Gong, Yifan Yao, Xinjing Huang, Jun Wang, Jianfeng Yu, Qi Guo, Yue Yu, Yan Zhang, Jin Wang, Hengtao Tao, Dasen Yan, Zexuan Yi, Fang Peng, Fangqing Jiang, Han Zhang, Lingfeng Deng, Yehong Zhang, Zhe Lin, Chao Zhang, Shaojie Zhang, Mingyue Guo, Shanzhi Gu, Gaojun Fan, Yaowei Wang, Xuefeng Jin, Qun Liu, and Yonghong Tian. 2021 · 2021
Closest in time.
CPM-2: large-scale cost-effective pre-trained language models
Zhengyan Zhang, Yuxian Gu, Xu Han, Shengqi Chen, Chaojun Xiao, Zhenbo Sun, Yuan Yao, Fanchao Qi, Jian Guan, Pei Ke, Yanzheng Cai, Guoyang Zeng, Zhixing Tan, Zhiyuan Liu, Minlie Huang, Wentao Han, Yang Liu, Xiaoyan Zhu, and Maosong Sun. 2021 · 2021
Closest in time.
Calibrate before use: Improving few-shot performance of language models
Tony Z. Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Closest in time.