Fetching the paper…
Reading the bibliography…
Transfer learning - i.e., further fine-tuning a pre-trained model on a downstream task - can confer significant advantages, including improved downstream performance, faster convergence, and better sample efficiency.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. M. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2019
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 1910
Earlier work this paper cites.
A comprehensive survey on transfer learning
F. Zhuang, Z. Qi, K. Duan, D. Xi, Y. Zhu, H. Zhu, H. Xiong, and Q. He · 1911
Earlier work this paper cites.
On the mathematical foundations of theoretical statistics
R. A. Fisher · 1922
Earlier work this paper cites.
Evaluating pruning methods
G. Thimm and E. Fiesler · 1995
Earlier work this paper cites.
Neural learning in structured parameter spaces - natural riemannian gradient
S. Amari · 1996
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
Y. LeCun · 1998
Earlier work this paper cites.
The pascal recognising textual entailment challenge
I. Dagan, O. Glickman, and B. Magnini · 2005
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
M. Roemmele, C. A. Bejan, and A. S. Gordon · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
J. Stallkamp, M. Schlipsing, J. Salmen, and C. Igel · 2011
Earlier work this paper cites.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Y. Li, J. Yosinski, J. Clune, H. Lipson, and J. Hopcroft · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Y. Yang, W.-t. Yih, and C. Meek · 2015
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization
C. D. Freeman and J. Bruna · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
S. Ruder · 2016
Earlier work this paper cites.
Sun database: Exploring a large collection of scene categories
J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, and A. Oliva · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
G. Cheng, J. Han, and X. Lu · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Essentially no barriers in neural network energy landscape
F. Draxler, K. Veschgini, M. Salmhofer, and F. Hamprecht · 2018
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
T. Garipov, P. Izmailov, D. Podoprikhin, D. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
J. Phang, T. Févry, and S. R. Bowman · 2018
Earlier work this paper cites.
Tackling the story ending biases in the story cloze test
R. Sharma, J. Allen, O. Bakhshandeh, and N. Mostafazadeh · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Cited alongside, same era.
Cosmos qa: Machine reading comprehension with contextual commonsense reasoning
L. Huang, R. Le Bras, C. Bhagavatula, and Y. Choi · 2019
Cited alongside, same era.
On the convergence of fedavg on non-iid data
X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang · 2019
Cited alongside, same era.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, and A. Peste · 2021
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, and W. Chen · 2021
Later among the works it cites.
Merging models with fisher-weighted averaging
M. Matena and C. Raffel · 2021
Later among the works it cites.
What to pre-train on? Efficient intermediate task selection
C. Poth, J. Pfeiffer, A. Rücklé, and I. Gurevych · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The CommitmentBank: Investigating projection in naturally occurring discourse
M.-C. d. Marneffe, M. Simons, and J. Tonhauser · 2019
Cited alongside, same era.
Adversarial NLI: A new benchmark for natural language understanding
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2019
Cited alongside, same era.
WiC: The word-in-context dataset for evaluating context-sensitive meaning representations
M. T. Pilehvar and J. Camacho-Collados · 2019
Cited alongside, same era.
Social iqa: Commonsense reasoning about social interactions
M. Sap, H. Rashkin, D. Chen, R. Le Bras, and Y. Choi · 2019
Cited alongside, same era.
Quartz: An open-domain dataset of qualitative relationship questions
O. Tafjord, M. Gardner, K. Lin, and P. Clark · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
PAWS: Paraphrase Adversaries from Word Scrambling
Y. Zhang, J. Baldridge, and L. He · 2019
Cited alongside, same era.
Later among the works it cites.
A survey on multi-task learning
Y. Zhang and Q. Yang · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries, 2022
S. K. Ainsworth, J. Hayase, and S. Srinivasa · 2022
Later among the works it cites.
PromptSource: An integrated development environment and repository for natural language prompts
S. H. Bach, V. Sanh, Z.-X. Yong, A. Webson, C. Raffel, N. V. Nayak, A. Sharma, T. Kim, M. S. Bari, T. Févry, et al · 2022
Later among the works it cites.
Fusing finetuned models for better pretraining, 2022
L. Choshen, E. Venezian, N. Slonim, and Y. Katz · 2022
Later among the works it cites.
Cold fusion: Collaborative descent for distributed multitask finetuning, 2022
S. Don-Yehiya, E. Venezian, C. Raffel, N. Slonim, Y. Katz, and L. Choshen · 2022
Later among the works it cites.
Patching open-vocabulary models by interpolating weights
G. Ilharco, M. Wortsman, S. Y. Gadre, S. Song, H. Hajishirzi, S. Kornblith, A. Farhadi, and L. Schmidt · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models, 2022
M. Li, S. Gururangan, T. Dettmers, M. Lewis, T. Althoff, N. A. Smith, and L. Zettlemoyer · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. A. Raffel · 2022
Later among the works it cites.
Model ratatouille: Recycling diverse models for out-of-distribution generalization
A. Ramé, K. Ahuja, J. Zhang, M. Cord, L. Bottou, and D. Lopez-Paz · 2022
Later among the works it cites.
Label sleuth: From unlabeled text to a classifier in a few hours
E. Shnarch, A. Halfon, A. Gera, M. Danilevsky, Y. Katsis, L. Choshen, M. S. Cooper, D. Epelboim, Z. Zhang, D. Wang, et al · 2022
Later among the works it cites.
Ul2: Unifying language learning paradigms
Y. Tay, M. Dehghani, V. Q. Tran, X. Garcia, J. Wei, X. Wang, H. W. Chung, D. Bahri, T. Schuster, S. Zheng, et al · 2022
Later among the works it cites.
When to use multi-task learning vs intermediate fine-tuning for pre-trained encoder transfer learning
O. Weller, K. Seppi, and M. Gardner · 2022
Later among the works it cites.
Improving few-shot generalization by exploring and exploiting auxiliary data
A. Albalak, C. Raffel, and W. Y. Wang · 2023
Closest in time.
Knowledge is a region in weight space for fine-tuned language models
A. Gueta, E. Venezian, C. Raffel, N. Slonim, Y. Katz, and L. Choshen · 2023
Closest in time.
Editing models with task arithmetic
G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi · 2023
Closest in time.
Dataless knowledge fusion by merging weights of language models
X. Jin, X. Ren, D. Preotiuc-Pietro, and P. Cheng · 2023
Closest in time.
REPAIR: REnormalizing permuted activations for interpolation repair
K. Jordan, H. Sedghi, O. Saukh, R. Entezari, and B. Neyshabur · 2023
Closest in time.
Editing implicit assumptions in text-to-image diffusion models
H. Orgad, B. Kawar, and Y. Belinkov · 2023
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
G. Ortiz-Jiménez, A. Favero, and P. Frossard · 2023
Closest in time.
Diverse weight averaging for out-of-distribution generalization
A. Ramé, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord · 2023
Closest in time.
An empirical study of multimodal model merging
Y.-L. Sung, L. Li, K. Lin, Z. Gan, M. Bansal, and L. Wang · 2023
Closest in time.
Exclusive supermask subnetwork training for continual learning
P. Yadav and M. Bansal · 2023
Closest in time.
Exploring continual learning for code generation models
P. Yadav, Q. Sun, H. Ding, X. Li, D. Zhang, M. Tan, P. Bhatia, X. Ma, R. Nallapati, M. K. Ramanathan, M. Bansal, and B. Xiang · 2023
Closest in time.