Fetching the paper…
Reading the bibliography…
Parameter-efficient fine-tuning (PEFT) for cross-task generalization consists in pre-training adapters on a multi-task training set before few-shot adaptation to test tasks.
Routing networks and the challenges of modular and compositional computation
C. Rosenbaum, I. Cases, M. Riemer, and T. Klinger · 1904
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Categorical reparameterization with Gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. P. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Modular networks: Learning to decompose neural computation
L. Kirsch, J. Kunze, and D. Barber · 2018
Earlier work this paper cites.
Systematic generalization: What is required and can it be learned?
D. Bahdanau, S. Murty, M. Noukhovitch, T. H. Nguyen, H. de Vries, and A. Courville · 2019
Earlier work this paper cites.
Modeling language variation and universals: A survey on typological linguistics for natural language processing
E. M. Ponti, H. O’Horan, Y. Berzak, I. Vulić, R. Reichart, T. Poibeau, E. Shutova, and A. Korhonen · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, et al · 2020
Earlier work this paper cites.
Exploring and predicting transferability across NLP tasks
T. Vu, T. Wang, T. Munkhdalai, A. Sordoni, A. Trischler, A. Mattarella-Micke, S. Maji, and M. Iyyer · 2020
Earlier work this paper cites.
MAD-G: Multilingual adapter generation for efficient cross-lingual transfer
A. Ansell, E. M. Ponti, J. Pfeiffer, S. Ruder, G. Glavaš, I. Vulić, and A. Korhonen · 2021
Cited alongside, same era.
Modular networks for compositional instruction following
R. Corona, D. Fried, C. Devin, D. Klein, and T. Darrell · 2021
Cited alongside, same era.
Recurrent independent mechanisms
A. Goyal, A. Lamb, J. Hoffmann, S. Sodhani, S. Levine, Y. Bengio, and B. Schölkopf · 2021
Cited alongside, same era.
Parameter-efficient multi-task fine-tuning for Transformers via shared hypernetworks
R. Karimi Mahabadi, S. Ruder, M. Dehghani, and J. Henderson · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Cited alongside, same era.
Continual learning via local module composition
O. Ostapenko, P. Rodriguez, M. Caccia, and L. Charlin · 2021
Composable sparse fine-tuning for cross-lingual transfer
A. Ansell, E. M. Ponti, A. Korhonen, and I. Vulić · 2022
Closest in time.
Ext5: Towards extreme multi-task scaling for transfer learning
V. Aribandi, Y. Tay, T. Schuster, J. Rao, H. S. Zheng, S. V. Mehta, H. Zhuang, V. Q. Tran, D. Bahri, J. Ni, J. Gupta, K. Hui, S. Ruder, and D. Metzler · 2022
Closest in time.
Attempt: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts
A. Asai, M. Salehi, M. E. Peters, and H. Hajishirzi · 2022
Closest in time.
IGLUE: A benchmark for transfer learning across modalities, tasks, and languages
E. Bugliarello, F. Liu, J. Pfeiffer, S. Reddy, D. Elliott, E. M. Ponti, and I. Vulić · 2022
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
AdapterFusion: Non-destructive task composition for transfer learning
J. Pfeiffer, A. Kamath, A. Rücklé, K. Cho, and I. Gurevych · 2021
Cited alongside, same era.
Inductive Bias and Modular Design for Sample-Efficient Neural Language Learning
E. Ponti · 2021
Cited alongside, same era.
Training neural networks with fixed sparse masks
Y.-L. Sung, V. Nair, and C. Raffel · 2021
Cited alongside, same era.
Spot: Better frozen model adaptation through soft prompt transfer
T. Vu, B. Lester, N. Constant, R. Al-Rfou, and D. Cer · 2021
Cited alongside, same era.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models
Z. Wang, Y. Tsvetkov, O. Firat, and Y. Cao · 2021
Cited alongside, same era.
CrossFit: A few-shot learning challenge for cross-task generalization in NLP
Q. Ye, B. Y. Lin, and X. Ren · 2021
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Closest in time.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning, 2022
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. Raffel · 2022
Closest in time.
Multitask prompted training enables zero-shot task generalization
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. L. Scao, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. V. Nayak, D. Datta, J. Chang, M. T. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Févry, J. A. Fries, R. Teehan, S. Biderman, L. Gao, T. Bers, T. Wolf, and A. M. Rush · 2022
Closest in time.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2022
Closest in time.
Adaptersoup: Weight averaging to improve generalization of pretrained language models
A. Chronopoulou, M. E. Peters, A. Fraser, and J. Dodge · 2023
Closest in time.
J. Pfeiffer, S. Ruder, I. Vulić, and E. M. Ponti · 2023
Closest in time.
Combining parameter-efficient modules for task-level generalisation
E. M. Ponti, A. Sordoni, Y. Bengio, and S. Reddy · 2023
Closest in time.
Compositional task representations for large language models
N. Shao, Z. Cai, H. xu, C. Liao, Y. Zheng, and Z. Yang · 2023
Closest in time.