Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Assessing bert’s syntactic abilities
Original
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Bert for joint intent classification and slot filling
Original
Qian Chen, Zhu Zhuo, and Wen Wang. 2019 · 1902
Earlier work this paper cites.
ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
Original
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2020 · 1904
Earlier work this paper cites.
Rare words: A major problem for contextualized embeddings and how to fix it by attentive mimicking
Original
Timo Schick and Hinrich Schütze. 2019 · 1904
Earlier work this paper cites.
Simple bert models for relation extraction and semantic role labeling
Original
Peng Shi and Jimmy Lin. 2019 · 1904
Earlier work this paper cites.
Towards better substitution-based word sense induction
Original
Asaf Amrami and Yoav Goldberg. 2019 · 1905
Earlier work this paper cites.
Adaptation of deep bidirectional multilingual transformers for russian language
Original
Yuri Kuratov and Mikhail Arkhipov. 2019 · 1905
Earlier work this paper cites.
How to fine-tune bert for text classification?
Original
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. 2020 · 1905
Earlier work this paper cites.
Enriching pre-trained language model with entity information for relation classification
Original
Shanchan Wu and Yifan He. 2019 · 1905
Earlier work this paper cites.
Comet: Commonsense transformers for automatic knowledge graph construction
Original
Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, and Yejin Choi. 2019 · 1906
Earlier work this paper cites.
UNIFIEDQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020 · 1907
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Does bert agree? evaluating knowledge of structure dependence through agreement relations
Original
Geoff Bacon and Terry Regier. 2019 · 1908
Earlier work this paper cites.
Bert for coreference resolution: Baselines and analysis
Original
Mandar Joshi, Omer Levy, Daniel S. Weld, and Luke Zettlemoyer. 2019 · 1908
Earlier work this paper cites.
Text summarization with pretrained encoders
Original
Yang Liu and Mirella Lapata. 2019 · 1908
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
Original
Joshua Feldman, Joe Davison, and Alexander M. Rush. 2019 · 1909
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Original
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 1909
Earlier work this paper cites.
Portuguese named entity recognition using bert-crf
Original
Fábio Souza, Rodrigo Nogueira, and Roberto Lotufo. 2020b · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Original
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020 · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Original
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 1911
Earlier work this paper cites.
Scalable zero-shot entity linking with dense entity retrieval
Original
Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettlemoyer. 2020a · 1911
Earlier work this paper cites.
Multi-domain dialogue state tracking as dynamic knowledge graph enhanced question answering
Original
Li Zhou and Kevin Small. 2020 · 1911
Earlier work this paper cites.
Zero-shot text classification with generative language models
Original
Raul Puri and Bryan Catanzaro. 2019 · 1912
Earlier work this paper cites.
olmpics – on what language model pre-training captures
Original
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 1912
Earlier work this paper cites.
Multilingual is not enough: Bert for finnish
Original
Antti Virtanen, Jenna Kanerva, Rami Ilo, Jouni Luoma, Juhani Luotolahti, Tapio Salakoski, Filip Ginter, and Sampo Pyysalo. 2019 · 1912
Earlier work this paper cites.
Bertje: A dutch bert model
Original
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. 2019 · 1912
Earlier work this paper cites.
Side-tuning: A baseline for network adaptation via additive side networks
Original
Jeffrey O Zhang, Alexander Sax, Amir Zamir, Leonidas Guibas, and Jitendra Malik. 2020b · 1912
Earlier work this paper cites.
Design of a knowledge-based report generator
Karen Kukich. 1983 · 1983
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A language modeling approach to information retrieval
Jay M. Ponte and W. Bruce Croft. 1998 · 1998
Earlier work this paper cites.
The CHILDES Project: Tools for analyzing talk. transcription format and programs , volume 1
Brian MacWhinney. 2000 · 2000
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Original
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Document language models, query models, and risk minimization for information retrieval
John Lafferty and Chengxiang Zhai. 2001 · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Learning cross-context entity representations from text
Original
Jeffrey Ling, Nicholas FitzGerald, Zifei Shan, Livio Baldini Soares, Thibault Févry, David Weiss, and Tom Kwiatkowski. 2020 · 2001
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Original
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020c · 2001
Earlier work this paper cites.
Tree-structured attention with hierarchical accumulation
Original
Xuan-Phi Nguyen, Shafiq Joty, Steven C. H. Hoi, and Richard Socher. 2020b · 2002
Earlier work this paper cites.
Incorporating bert into neural machine translation
Original
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020 · 2002
Earlier work this paper cites.
Dare: Data augmented relation extraction with gpt-2
Original
Yannis Papanikolaou and Andrea Pierleoni. 2020 · 2004
Earlier work this paper cites.
A simple language model for task-oriented dialogue
Original
Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020 · 2005
Earlier work this paper cites.
Muss: Multilingual unsupervised sentence simplification by mining paraphrases
Original
Louis Martin, Angela Fan, Éric de la Clergerie, Antoine Bordes, and Benoît Sagot. 2021 · 2005
Earlier work this paper cites.
Evaluating German transformer language models with syntactic agreement tests
Original
Karolina Zaczynska, Nils Feldhus, Robert Schwarzenberg, Aleksandra Gabryszak, and Sebastian Möller. 2020 · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
On data augmentation for extreme multi-label classification
Original
Danqing Zhang, Tao Li, Haiyang Zhang, and Bing Yin. 2020a · 2009
Earlier work this paper cites.
The turking test: Can language models understand instructions?
Original
Avia Efrat and Omer Levy. 2020 · 2010
Earlier work this paper cites.
Why Does Unsupervised Pre-training Help Deep Learning?
Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio. 2010 · 2010
Earlier work this paper cites.
Probing and fine-tuning reading comprehension models for few-shot event extraction
Original
Rui Feng, Jie Yuan, and Chao Zhang. 2020a · 2010
Earlier work this paper cites.
Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs
Original
Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi. 2020 · 2010
Earlier work this paper cites.
Automated concatenation of embeddings for structured prediction
Original
Xinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, and Kewei Tu. 2021c · 2010
Earlier work this paper cites.
Strongly incremental constituency parsing with graph neural networks
Original
Kaiyu Yang and Jia Deng. 2020 · 2010
Earlier work this paper cites.
A survey of knowledge-enhanced text generation
Original
Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang. 2021b · 2010
Earlier work this paper cites.
Paired representation learning for event and entity coreference
Original
Xiaodong Yu, Wenpeng Yin, and Dan Roth. 2020 · 2010
Earlier work this paper cites.
Language model is all you need: Natural language understanding as question answering
Original
Mahdi Namazifar, Alexandros Papangelis, Gokhan Tur, and Dilek Hakkani-Tür. 2020 · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Original
Zeyuan Allen-Zhu and Yuanzhi Li. 2021 · 2012
Earlier work this paper cites.
Parameter-Efficient Transfer Learning with Diff Pruning
Original
Demi Guo, Alexander M. Rush, and Yoon Kim. 2021 · 2012
Earlier work this paper cites.
Who did what to whom? a contrastive study of syntacto-semantic dependencies
Angelina Ivanova, Stephan Oepen, Lilja Øvrelid, and Dan Flickinger. 2012 · 2012
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Xlm-t: Scaling up multilingual machine translation with pretrained cross-lingual transformer encoders
Original
Shuming Ma, Jian Yang, Haoyang Huang, Zewen Chi, Li Dong, Dongdong Zhang, Hany Hassan Awadalla, Alexandre Muzio, Akiko Eriguchi, Saksham Singhal, Xia Song, Arul Menezes, and Furu Wei. 2020 · 2012
Earlier work this paper cites.
Few-shot text generation with pattern-exploiting training
Original
Timo Schick and Hinrich Schütze. 2020 · 2012
Earlier work this paper cites.
Universal Conceptual Cognitive Annotation (UCCA)
Omri Abend and Ari Rappoport. 2013 · 2013
Earlier work this paper cites.
Abstract Meaning Representation for sembanking
Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013 · 2013
Earlier work this paper cites.
Recognizing textual entailment: Models and applications
Ido Dagan, Dan Roth, Mark Sammons, and Fabio Massimo Zanzotto. 2013 · 2013
Earlier work this paper cites.
Constructing information networks using one single model
Qi Li, Heng Ji, Yu Hong, and Sujian Li. 2014 · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Earlier work this paper cites.
What makes ImageNet good for transfer learning?
Original
Minyoung Huh, Pulkit Agrawal, and Alexei A. Efros. 2016 · 2016
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability
Been Kim, Oluwasanmi Koyejo, and Rajiv Khanna. 2016 · 2016
Earlier work this paper cites.
End-to-end relation extraction using LSTMs on sequences and tree structures
Makoto Miwa and Mohit Bansal. 2016 · 2016
Earlier work this paper cites.
Joint event extraction via recurrent neural networks
Thien Huu Nguyen, Kyunghyun Cho, and Ralph Grishman. 2016 · 2016
Earlier work this paper cites.
”why should I trust you?”: Explaining the predictions of any classifier
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Reading Wikipedia to answer open-domain questions
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Incidental supervision: Moving beyond supervised learning
Dan Roth. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin J. Wainwright, and Michael I. Jordan. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Colorless green recurrent networks dream hierarchically
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Jointly predicting predicates and arguments in neural semantic role labeling
Luheng He, Kenton Lee, Omer Levy, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John F. Canny, and Zeynep Akata. 2018 · 2018
Earlier work this paper cites.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Earlier work this paper cites.