Fetching the paper…
Reading the bibliography…
Training large foundation models using self-supervised objectives on unlabeled data, followed by fine-tuning on downstream tasks, has emerged as a standard procedure.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng et al., · 2009
Earlier work this paper cites.
“Microsoft coco: Common objects in context,”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, · 2014
Earlier work this paper cites.
“VQA: Visual question answering,”
Stanislaw Antol et al., · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, et al., · 2017
Earlier work this paper cites.
“The lj speech dataset,”
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke et al., · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,” 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“Generalized sliced wasserstein distances,”
Soheil Kolouri et al., · 2019
Earlier work this paper cites.
“A unified framework for domain adaptation using metric learning on manifolds,”
Sridhar Mahadevan, Bamdev Mishra, and Shalini Ghosh, · 2019
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald, · 2019
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Earlier work this paper cites.
“Parameter-efficient transfer learning for nlp,”
Neil Houlsby et al., · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown et al., · 2020
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy et al., · 2020
Cited alongside, same era.
“Model fusion via optimal transport,”
Sidak Pal Singh and Martin Jaggi, · 2020
Cited alongside, same era.
“Class-incremental learning via deep model consolidation,”
Junting Zhang, Jie Zhang, Shalini Ghosh, Dawei Li, Serafettin Tasci, Larry P. Heck, Heming Zhang, and C.-C. Jay Kuo, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
J. Kahn et al., · 2020
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
“Beats: Audio pre-training with acoustic tokenizers,”
Sanyuan Chen et al., · 2022
Later among the works it cites.
“lo-fi: distributed fine-tuning without communication,”
Mitchell Wortsman et al., · 2022
Later among the works it cites.
“Fusing finetuned models for better pretraining,”
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz, · 2022
Later among the works it cites.
“Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,”
Mitchell Wortsman et al., · 2022
Later among the works it cites.
“Otkge: Multi-modal knowledge graph embeddings via optimal transport,”
Zongsheng Cao et al., · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Pretrained transformers as universal computation engines,”
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch, · 2021
Cited alongside, same era.
“Voice2series: Reprogramming acoustic models for time series classification,”
Chao-Han Huck Yang, Yun-Yun Tsai, and Pin-Yu Chen, · 2021
Cited alongside, same era.
“Multimodal neurons in artificial neural networks,”
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah, · 2021
Cited alongside, same era.
“I can’t believe there’s no images! learning visual tasks using only language data,”
Sophia Gu, Christopher Clark, and Aniruddha Kembhavi, · 2022
Cited alongside, same era.
“Can wikipedia help offline reinforcement learning?,”
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu, · 2022
Cited alongside, same era.
“Pretraining with artificial language: Studying transferable knowledge in language models,”
Ryokan Ri and Yoshimasa Tsuruoka, · 2022
Cited alongside, same era.
“Merging models with fisher-weighted averaging,”
Michael S Matena and Colin A Raffel, · 2022
Later among the works it cites.
“Editing models with task arithmetic,”
Gabriel Ilharco et al., · 2022
Later among the works it cites.
“Are all layers created equal?,”
Chiyuan Zhang, Samy Bengio, and Yoram Singer, · 2022
Later among the works it cites.
“Lora: Low-rank adaptation of large language models,”
Edward J Hu et al., · 2022
Later among the works it cites.
“Multimodal neurons in pretrained text-only transformers,”
Sarah Schwettmann, Neil Chowdhury, and Antonio Torralba, · 2023
Closest in time.
“An Empirical Study of Multimodal Model Merging,” Apr. 2023,
Yi-Lin Sung, Linjie Li, Kevin Lin, Zhe Gan, Mohit Bansal, and Lijuan Wang, · 2023
Closest in time.
“Adaptersoup: Weight averaging to improve generalization of pretrained language models,”
Alexandra Chronopoulou, Matthew E Peters, Alexander Fraser, and Jesse Dodge, · 2023
Closest in time.
“How to estimate model transferability of pre-trained speech models?,”
Zih-Ching Chen et al., · 2023
Closest in time.