Fetching the paper…
Reading the bibliography…
Plug-and-play functionality allows deep learning models to adapt well to different tasks without requiring any parameters modified.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 1902
Earlier work this paper cites.
XLNet: Generalized Autoregressive Pretraining for Language Understanding
Yang, Z.; Dai, Z.; Yang, Y.; Carbonell, J.; Salakhutdinov, R.; and Le, Q. V. 2020 · 1906
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Side-Tuning: A Baseline for Network Adaptation via Additive Side Networks
Zhang, J. O.; Sax, A.; Zamir, A.; Guibas, L.; and Malik, J. 2020 · 1912
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, M. P.; Santorini, B.; and Marcinkiewicz, M. A. 1993 · 1993
Earlier work this paper cites.
Introduction to the CoNLL-2000 Shared Task Chunking
Tjong Kim Sang, E. F.; and Buchholz, S. 2000 · 2000
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
Lafferty, J.; McCallum, A.; and Pereira, F. C. 2001 · 2001
Earlier work this paper cites.
Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference
Schick, T.; and Schütze, H. 2021 · 2001
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Sang, E. F.; and De Meulder, F. 2003 · 2003
Earlier work this paper cites.
Masking as an Efficient Alternative to Finetuning for Pretrained Language Models
Zhao, M.; Lin, T.; Mi, F.; Jaggi, M.; and Schütze, H. 2020 · 2004
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V.; Wolf, T.; and Rush, A. M. 2020 · 2005
Earlier work this paper cites.
Plug-and-Play Conversational Models
Madotto, A.; Ishii, E.; Lin, Z.; Dathathri, S.; and Fung, P. 2020 · 2010
Earlier work this paper cites.
Parameter-Efficient Transfer Learning with Diff Pruning
Guo, D.; Rush, A. M.; and Kim, Y. 2021 · 2012
Earlier work this paper cites.
Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation
Cho, K.; van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Conditional random field with high-order dependencies for sequence labeling and segmentation
Cuong, N. V.; Ye, N.; Lee, W. S.; and Chieu, H. L. 2014 · 2014
Earlier work this paper cites.
Learning character-level representations for part-of-speech tagging
Dos Santos, C.; and Zadrozny, B. 2014 · 2014
Cited alongside, same era.
Lexicon infused phrase embeddings for named entity resolution
Passos, A.; Kumar, V.; and McCallum, A. 2014 · 2014
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Cited alongside, same era.
Named entity recognition with bidirectional LSTM-CNNs
Chiu, J. P.; and Nichols, E. 2016 · 2016
Cited alongside, same era.
End-to-end sequence labeling via bi-directional lstm-cnns-crf
Ma, X.; and Hovy, E. 2016 · 2016
Cited alongside, same era.
Attention is all you need
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020 · 2020
Later among the works it cites.
Hierarchical contextualized representation for named entity recognition
Luo, Y.; Xiao, F.; and Zhao, H. 2020 · 2020
Later among the works it cites.
AdapterHub: A Framework for Adapting Transformers
Pfeiffer, J.; Rücklé, A.; Poth, C.; Kamath, A.; Vulić, I.; Ruder, S.; Cho, K.; and Gurevych, I. 2020 · 2020
Later among the works it cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Later among the works it cites.
KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction
Chen, X.; Zhang, N.; Xie, X.; Deng, S.; Yao, Y.; Tan, C.; Huang, F.; Si, L.; and Chen, H. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Cited alongside, same era.
Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network
Sharma, H.; Park, J.; Suda, N.; Lai, L.; Chau, B.; Chandra, V.; and Esmaeilzadeh, H. 2018 · 2018
Cited alongside, same era.
Context is Key: Grammatical Error Detection with Contextual Word Representations
Bell, S.; Yannakoudakis, H.; and Rei, M. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J.; and Carbin, M. 2019 · 2019
Cited alongside, same era.
Closest in time.
Template-Based Named Entity Recognition Using BART
Cui, L.; Wu, Y.; Liu, J.; Yang, S.; and Zhang, Y. 2021 · 2021
Closest in time.
WARP: Word-level Adversarial ReProgramming
Hambardzumyan, K.; Khachatrian, H.; and May, J. 2021 · 2021
Closest in time.
PTR: Prompt Tuning with Rules for Text Classification
Han, X.; Zhao, W.; Ding, N.; Liu, Z.; and Sun, M. 2021 · 2021
Closest in time.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Closest in time.
Liu, X.; Zheng, Y.; Du, Z.; Ding, M.; Qian, Y.; Yang, Z.; and Tang, J. 2021 · 2021
Closest in time.
Hardware Acceleration of Fully Quantized BERT for Efficient Natural Language Processing
Liu, Z.; Li, G.; and Cheng, J. 2021 · 2021
Closest in time.
Entailment as Few-Shot Learner
Wang, S.; Fang, H.; Khabsa, M.; Mao, H.; and Ma, H. 2021 · 2021
Closest in time.
A Unified Generative Framework for Various NER Subtasks
Yan, H.; Gui, T.; Dai, J.; Guo, Q.; Zhang, Z.; and Qiu, X. 2021 · 2021
Closest in time.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Zaken, E. B.; Ravfogel, S.; and Goldberg, Y. 2021 · 2021
Closest in time.