Fetching the paper…
Reading the bibliography…
While fine-tuning pretrained models has become common practice, these models often underperform outside their specific domains.
On the mathematical foundations of theoretical statistics
R. A. Fisher · 1922
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
B. W. Matthews · 1975
Earlier work this paper cites.
Evidence for universality and cultural variation of differential emotion response patterning
K. R. Scherer and H. G. Wallbott · 1994
Earlier work this paper cites.
Adapting arbitrary normal mutation distributions in evolution strategies: The covariance matrix adaptation
N. Hansen and A. Ostermeier · 1996
Earlier work this paper cites.
The mnist database of handwritten digits, 1998
Y. LeCun · 1998
Earlier work this paper cites.
Emotions from text: machine learning for text-based emotion prediction
C. O. Alm, D. Roth, and R. Sproat · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
B. Dolan and C. Brockett · 2005
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
D. Giampiccolo, B. Magnini, I. Dagan, and W. B. Dolan · 2007
Earlier work this paper cites.
Semeval-2007 task 14: Affective text
C. Strapparava and R. Mihalcea · 2007
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng, et al · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
M. Roemmele, C. A. Bejan, and A. S. Gordon · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
J. Stallkamp, M. Schlipsing, J. Salmen, and C. Igel · 2011
Earlier work this paper cites.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern · 2012
Earlier work this paper cites.
#emotional tweets
S. M. Mohammad · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, and C. Potts · 2013
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Y. Li, J. Yosinski, J. Clune, H. Lipson, and J. Hopcroft · 2015
Earlier work this paper cites.
Sentiment, emotion, purpose, and style in electoral tweets
S. M. Mohammad, X. Zhu, S. Kiritchenko, and J. Martin · 2015
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering
Y. Yang, W.-t. Yih, and C. Meek · 2015
Earlier work this paper cites.
Sun database: Exploring a large collection of scene categories
J. Xiao, K. A. Ehinger, J. Hays, A. Torralba, and A. Oliva · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
G. Cheng, J. Han, and X. Lu · 2017
Earlier work this paper cites.
Dailydialog: A manually labelled multi-turn dialogue dataset
Y. Li, H. Su, X. Shen, W. Li, Z. Cao, and S. Niu · 2017
Earlier work this paper cites.
Grounded emotions
V. Liu, C. Banea, and R. Mihalcea · 2017
Earlier work this paper cites.
Wassa-2017 shared task on emotion intensity
S. Mohammad and F. Bravo-Marquez · 2017
Earlier work this paper cites.
Annotation, modelling and analysis of fine-grained emotions on a stance and sentiment detection corpus
H. Schuff, J. Barnes, J. Mohme, S. Padó, and R. Klinger · 2017
Earlier work this paper cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
An analysis of annotated corpora for emotion classification in text
L. A. M. Oberländer and R. Klinger · 2018
Earlier work this paper cites.
Ensemble learning: A survey
O. Sagi and L. Rokach · 2018
Earlier work this paper cites.
Tackling the story ending biases in the story cloze test
R. Sharma, J. Allen, O. Bakhshandeh, and N. Mostafazadeh · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Cited alongside, same era.
Cosmos qa: Machine reading comprehension with contextual commonsense reasoning
L. Huang, R. Le Bras, C. Bhagavatula, and Y. Choi · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. D. M.-W. C. Kenton and L. K. Toutanova · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Later among the works it cites.
Ensemble of averages: Improving model selection and boosting performance in domain generalization
D. Arpit, H. Wang, Y. Zhou, and C. Xiong · 2022
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
S. H. Bach, V. Sanh, Z.-X. Yong, A. Webson, C. Raffel, N. V. Nayak, A. Sharma, T. Kim, M. S. Bari, T. Fevry, et al · 2022
Later among the works it cites.
Fusing finetuned models for better pretraining
L. Choshen, E. Venezian, N. Slonim, and Y. Katz · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The CommitmentBank: Investigating projection in naturally occurring discourse
M.-C. d. Marneffe, M. Simons, and J. Tonhauser · 2019
Cited alongside, same era.
WiC: The word-in-context dataset for evaluating context-sensitive meaning representations
M. T. Pilehvar and J. Camacho-Collados · 2019
Cited alongside, same era.
Social iqa: Commonsense reasoning about social interactions
M. Sap, H. Rashkin, D. Chen, R. Le Bras, and Y. Choi · 2019
Cited alongside, same era.
Quartz: An open-domain dataset of qualitative relationship questions
O. Tafjord, M. Gardner, K. Lin, and P. Clark · 2019
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Cited alongside, same era.
Neural network acceptability judgments
A. Warstadt, A. Singh, and S. Bowman · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Cited alongside, same era.
Patching open-vocabulary models by interpolating weights
G. Ilharco, M. Wortsman, S. Y. Gadre, S. Song, H. Hajishirzi, S. Kornblith, A. Farhadi, and L. Schmidt · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, and C. A. Raffel · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
M. S. Matena and C. A. Raffel · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
A. Rame, M. Kirchmeyer, T. Rahier, A. Rakotomamonjy, P. Gallinari, and M. Cord · 2022
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. L. Scao, A. Raja, et al · 2022
Later among the works it cites.
Label sleuth: From unlabeled text to a classifier in a few hours
E. Shnarch, A. Halfon, A. Gera, M. Danilevsky, Y. Katsis, L. Choshen, M. S. Cooper, D. Epelboim, Z. Zhang, D. Wang, et al · 2022
Later among the works it cites.
Black-box tuning for language-model-as-a-service
T. Sun, Y. Shao, H. Qian, X. Huang, and X. Qiu · 2022
Later among the works it cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, et al · 2022
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
S. K. Ainsworth, J. Hayase, and S. Srinivasa · 2023
Later among the works it cites.
Lorahub: Efficient cross-task generalization via dynamic lora composition
C. Huang, Q. Liu, B. Y. Lin, T. Pang, C. Du, and M. Lin · 2023
Later among the works it cites.
Editing models with task arithmetic
G. Ilharco, M. T. Ribeiro, M. Wortsman, S. Gururangan, L. Schmidt, H. Hajishirzi, and A. Farhadi · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
X. Jin, X. Ren, D. Preotiuc-Pietro, and P. Cheng · 2023
Later among the works it cites.
Editing implicit assumptions in text-to-image diffusion models
H. Orgad, B. Kawar, and Y. Belinkov · 2023
Later among the works it cites.
Code llama: Open foundation models for code
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. Défossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operation
J. Zhang, J. Liu, J. He, et al · 2023
Later among the works it cites.
Proving linear mode connectivity of neural networks via optimal transport
D. Ferbach, B. Goujaud, G. Gidel, and A. Dieuleveut · 2024
Closest in time.
Cade: Cosine annealing differential evolution for spiking neural network
R. Jiang, G. Du, S. Yu, Y. Guo, S. K. Goh, and H.-K. Tang · 2024
Closest in time.
No train but gain: Language arithmetic for training-free language adapters enhancement
M. Klimaszewski, P. Andruszkiewicz, and A. Birch · 2024
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
G. Ortiz-Jimenez, A. Favero, and P. Frossard · 2024
Closest in time.
Zipit! merging models from different tasks without training
G. Stoica, D. Bolya, J. Bjorner, P. Ramesh, T. Hearn, and J. Hoffman · 2024
Closest in time.
Knowledge fusion of large language models
F. Wan, X. Huang, D. Cai, X. Quan, W. Bi, and S. Shi · 2024
Closest in time.
Training-free pretrained model merging
Z. Xu, K. Yuan, H. Wang, Y. Wang, M. Song, and J. Song · 2024
Closest in time.
Ties-merging: Resolving interference when merging models
P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal · 2024
Closest in time.
Going beyond linear mode connectivity: The layerwise linear feature connectivity
Z. Zhou, Y. Yang, X. Yang, J. Yan, and W. Hu · 2024
Closest in time.