Fetching the paper…
Reading the bibliography…
Prompt Transfer (PoT) is a recently-proposed approach to improve prompt-tuning, by initializing the target prompt with the existing prompt trained on similar source tasks.
C. Spearman, “The proof and measurement of association between two things.” American Journal of Psychology , 1904
1904
Earlier work this paper cites.
E. T. K. Sang and F. De Meulder, “Introduction to the conll-2003 shared task: Language-independent named entity recognition,” in NAACL , 2003
2003
Earlier work this paper cites.
X. Carreras and L. Màrquez, “Introduction to the CoNLL-2004 shared task: Semantic role labeling,” in NAACL , 2004
2004
Earlier work this paper cites.
X. Carreras and L. Màrquez, “Introduction to the conll-2005 shared task: Semantic role labeling,” in CoNLL-2005 , 2005
2005
Earlier work this paper cites.
E. Hovy, M. Marcus, M. Palmer, L. Ramshaw, and R. Weischedel, “Ontonotes: the 90% solution,” in NAACL , 2006
2006
Earlier work this paper cites.
S. Pradhan, A. Moschitti, N. Xue, O. Uryupina, and Y. Zhang, “CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes,” in EMNLP , 2012
2012
Earlier work this paper cites.
G. Hinton, O. Vinyals, J. Dean et al. , “Distilling the knowledge in a neural network,” in NeurIPS , 2015
2015
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100, 000+ questions for machine comprehension of text,” in EMNLP , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
T. Furlanello, Z. Lipton, M. Tschannen, L. Itti, and A. Anandkumar, “Born again neural networks,” in ICML , 2018
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in EMNLP , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” arXiv , 2019
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in ICML , 2019
2019
Earlier work this paper cites.
S. Hahn and H. Choi, “Self-knowledge distillation in natural language processing,” in RANLP , 2019
2019
Earlier work this paper cites.
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Superglue: A stickier benchmark for general-purpose language understanding systems,” in NeurIPS , 2019
2019
Earlier work this paper cites.
P. He, X. Liu, J. Gao, and W. Chen, “Deberta: Decoding-enhanced bert with disentangled attention,” in ICLR , 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” JMLR , 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
F. Yuan, X. He, A. Karatzoglou, and L. Zhang, “Parameter-efficient transfer from sequential behaviors for user modeling and recommendation,” in SIGIR , 2020
2020
Earlier work this paper cites.
Y. Pruksachatkun, J. Phang, H. Liu, P. M. Htut, X. Zhang, R. Y. Pang, C. Vania, K. Kann, and S. Bowman, “Intermediate-task transfer learning with pretrained language models: When and why does it work?” in ACL , 2020
2020
Earlier work this paper cites.
S. Chen, Y. Hou, Y. Cui, W. Che, T. Liu, and X. Yu, “Recall and learn: Fine-tuning deep pretrained language models with less forgetting,” in EMNLP , 2020
2020
Cited alongside, same era.
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “Spanbert: Improving pre-training by representing and predicting spans,” TACL , 2020
2020
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL , 2020
2020
Cited alongside, same era.
R. Guan, H. Zhang, Y. Liang, F. Giunchiglia, L. Huang, and X. Feng, “Deep feature-based text clustering and its explanation,” IEEE Transactions on Knowledge and Data Engineering , 2020
2020
Cited alongside, same era.
T. Vu, B. Lester, N. Constant, R. Al-Rfou, and D. Cer, “SPoT: Better frozen model adaptation through soft prompt transfer,” in ACL , 2022
2022
Closest in time.
Y. Su, X. Wang, Y. Qin, C.-M. Chan, Y. Lin, H. Wang, K. Wen, Z. Liu, P. Li, J. Li et al. , “On transferability of prompt tuning for natural language processing,” in NAACL , 2022
2022
Closest in time.
Q. Zhong, L. Ding, Y. Zhan, Y. Qiao, Y. Wen, L. Shen, J. Liu, B. Yu, B. Du, Y. Chen et al. , “Toward efficient language model pretraining and downstream adaptation via self-evolution: A case study on superglue,” arXiv , 2022
2022
Closest in time.
Q. Zhong, L. Ding, L. Shen, P. Mi, J. Liu, B. Du, and D. Tao, “Improving sharpness-aware minimization with fisher mask for better generalization on language models,” in Findings of EMNLP , 2022
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Li, B. Chiu, S. Feng, and H. Wang, “Few-shot named entity recognition via meta-learning,” IEEE Transactions on Knowledge and Data Engineering , 2020
2020
Cited alongside, same era.
J. Li, A. Sun, and Y. Ma, “Neural named entity boundary detection,” IEEE Transactions on Knowledge and Data Engineering , 2020
2020
Cited alongside, same era.
J. Li, A. Sun, J. Han, and C. Li, “A survey on deep learning for named entity recognition,” IEEE Transactions on Knowledge and Data Engineering , 2020
2020
Cited alongside, same era.
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh, “Autoprompt: Eliciting knowledge from language models with automatically generated prompts,” in EMNLP , 2020
2020
Cited alongside, same era.
G. Xu, Z. Liu, X. Li, and C. C. Loy, “Knowledge distillation meets self-supervision,” in ECCV , 2020
2020
Cited alongside, same era.
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t stop pretraining: Adapt language models to domains and tasks,” in ACL , 2020
2020
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in EMNLP , 2021
2021
Cited alongside, same era.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in ACL , 2021
2021
Cited alongside, same era.
X. Liu, K. Ji, Y. Fu, W. Tam, Z. Du, Z. Yang, and J. Tang, “P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks,” in ACL , 2022
2022
Closest in time.
Y. Gu, X. Han, Z. Liu, and M. Huang, “PPT: Pre-trained prompt tuning for few-shot learning,” in ACL , 2022
2022
Closest in time.
X. Han, W. Zhao, N. Ding, Z. Liu, and M. Sun, “Ptr: Prompt tuning with rules for text classification,” AI Open , 2022
2022
Closest in time.
A. Asai, M. Salehi, M. E. Peters, and H. Hajishirzi, “Attempt: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts,” in EMNLP , 2022
2022
Closest in time.
T. Jiang, J. Jiao, S. Huang, Z. Zhang, D. Wang, F. Zhuang, F. Wei, H. Huang, D. Deng, and Q. Zhang, “Promptbert: Improving bert sentence embeddings with prompts,” in EMNLP , 2022
2022
Closest in time.
B. Wang, L. Ding, Q. Zhong, X. Li, and D. Tao, “A contrastive cross-channel data augmentation framework for aspect-based sentiment analysis,” in COLING , 2022
2022
Closest in time.
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin et al. , “Opt: Open pre-trained transformer language models,” arXiv , 2022
2022
Closest in time.
Q. Zhong, L. Ding, K. Peng, J. Liu, B. Du, L. Shen, Y. Zhan, and D. Tao, “Bag of tricks for effective language model pretraining and downstream adaptation: A case study on glue,” arXiv , 2023
2023
Closest in time.
Q. Zhong, L. Ding, J. Liu, B. Du, and D. Tao, “Self-evolution learning for discriminative language model pretraining,” in Findings of ACL , 2023
2023
Closest in time.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , 2023
2023
Closest in time.
Q. Zhong, L. Ding, J. Liu, B. Du, and D. Tao, “E2s2: Encoding-enhanced sequence-to-sequence pretraining for language understanding and generation,” IEEE Transactions on Knowledge and Data Engineering , 2023
2023
Closest in time.
Q. Zhong, L. Ding, J. Liu, B. Du, H. Jin, and D. Tao, “Knowledge graph augmented network towards multiview representation learning for aspect-based sentiment analysis,” IEEE Transactions on Knowledge and Data Engineering , 2023
2023
Closest in time.
Q. Zhong, L. Ding, J. Liu, X. Liu, M. Zhang, B. Du, and D. Tao, “Revisiting token dropping strategy in efficient bert pretraining,” in ACL , 2023
2023
Closest in time.
H. Huang, X. Liu, G. Shi, and Q. Liu, “Event extraction with dynamic prefix tuning and relevance retrieval,” IEEE Transactions on Knowledge and Data Engineering , 2023
2023
Closest in time.
X. Peng, C. Xing, P. K. Choubey, C.-S. Wu, and C. Xiong, “Model ensemble instead of prompt fusion: a sample-specific knowledge transfer method for few-shot prompt tuning,” in ICLR , 2023
2023
Closest in time.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al. , “Llama 2: Open foundation and fine-tuned chat models,” arXiv , 2023
2023
Closest in time.
Q. Zhong, L. Ding, L. Shen, J. Liu, B. Du, and D. Tao, “Revisiting knowledge distillation for autoregressive language models,” arXiv , 2024
2024
Closest in time.