Fetching the paper…
Reading the bibliography…
Prompt-Tuning is a new paradigm for finetuning pre-trained language models in a parameter-efficient way.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Hypernetworks
Ha, D., Dai, A. M., and Le, Q. V · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
Augmenting self-attention with persistent memory
Sukhbaatar, S., Grave, E., Lample, G., Jegou, H., and Joulin, A · 2019
Earlier work this paper cites.
Continual learning with hypernetworks
von Oswald, J., Henning, C., Sacramento, J., and Grewe, B. F · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Cited alongside, same era.
Dehghani, M., Arnab, A., Beyer, L., Vaswani, A., and Tay, Y · 2021
Later among the works it cites.
WARP: Word-level Adversarial ReProgramming
Hambardzumyan, K., Khachatrian, H., and May, J · 2021
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning, 2021
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2021
Later among the works it cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Karimi Mahabadi, R., Ruder, S., Dehghani, M., and Henderson, J · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Cited alongside, same era.
Hypergrid: Efficient multi-task transformers with grid-wise decomposable hyper projections
Tay, Y., Zhao, Z., Bahri, D., Metzler, D., and Juan, D.-C · 2020
Cited alongside, same era.
Small towers make big differences, 2020
Wang, Y., Zhao, Z., Dai, B., Fifty, C., Lin, D., Hong, L., and Chi, E. H · 2020
Cited alongside, same era.
Understanding and improving information transfer in multi-task learning
Wu, S., Zhang, H. R., and Ré, C · 2020
Cited alongside, same era.
Muppet: Massive multi-task representations with pre-finetuning
Aghajanyan, A., Gupta, A., Shrivastava, A., Chen, X., Zettlemoyer, L., and Gupta, S · 2021
Cited alongside, same era.
Ext5: Towards extreme multi-task scaling for transfer learning, 2021
Aribandi, V., Tay, Y., Schuster, T., Rao, J., Zheng, H. S., Mehta, S. V., Zhuang, H., Tran, V. Q., Bahri, D., Ni, J., Gupta, J., Hui, K., Ruder, S., and Metzler, D · 2021
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S
Cited in the paper.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S
Cited in the paper.
Li, X. L. and Liang, P · 2021
Later among the works it cites.
Gpt understands, too, 2021
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization, 2021
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Biderman, S., Gao, L., Bers, T., Wolf, T., and Rush, A. M · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners, 2021
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Ravfogel, S., and Goldberg, Y · 2021
Later among the works it cites.