Fetching the paper…
Reading the bibliography…
Fine-tuning pre-trained models provides significant advantages in downstream performance.
Microsoft Research Paraphrase Corpus
Dolan, William B and Brockett, Chris. 2005 · 2005
Earlier work this paper cites.
Recognizing Textual Entailment Dataset
Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel. 2009 · 2009
Earlier work this paper cites.
Stanford Sentiment Treebank
Socher, Richard and Perelygin, Alex and Wu, Jean and Chuang, Jason and Manning, Christopher D and Ng, Andrew and Potts, Christopher. 2013 · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft. 2016 · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H. B. McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2016 · 2016
Earlier work this paper cites.
Stanford Question Answering Dataset
Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Quora Question Pairs
Shankar Iyer and Nikhil Dandekar and and Kornél Csernai. 2017 · 2017
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Multi-Genre Natural Language Inference (MultiNLI) Corpus
Williams, Adina and Nangia, Nikita and Bowman, Samuel R. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
On the convergence of fedavg on non-iid data
Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. 2020 · 2020
Cited alongside, same era.
Pre-trained models for natural language processing: A survey
XiPeng Qiu, TianXiang Sun, YiGe Xu, YunFan Shao, Ning Dai, and XuanJing Huang. 2020 · 2020
Cited alongside, same era.
Optimizing mode connectivity via neuron alignment
N. Joseph Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk, Prasanna Sattigeri, and Rongjie Lai. 2020 · 2020
Cited alongside, same era.
Fusing finetuned models for better pretraining
Leshem Choshen, Elad Venezian, Noam Slonim, and Yoav Katz. 2022 · 2022
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur. 2022 · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami. 2022 · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Michael Matena and Colin Raffel. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. 2021 · 2021
Cited alongside, same era.
Efficiently identifying task groupings for multi-task learning
Chris Fifty, Ehsan Amid, Zhe Zhao, Tianhe Yu, Rohan Anil, and Chelsea Finn. 2021 · 2021
Cited alongside, same era.
Pre-trained models: Past, present and future
Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Yuan Yao, Ao Zhang, Liang Zhang, Wentao Han, Minlie Huang, Qin Jin, Yanyan Lan, Yang Liu, Zhiyuan Liu, Zhiwu Lu, Xipeng Qiu, Ruihua Song, Jie Tang, Ji-Rong Wen, Jinhui Yuan, Wayne Xin Zhao, and Jun Zhu. 2021 · 2021
Cited alongside, same era.
Measuring data leakage in machine-learning models with fisher information
Awni Hannun, Chuan Guo, and Laurens van der Maaten. 2021 · 2021
Cited alongside, same era.
Group fisher pruning for practical network compression
Liyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou, Jing-Hao Xue, Xinjiang Wang, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang. 2021 · 2021
Cited alongside, same era.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heinz, and Dan Roth. 2021 · 2021
Cited alongside, same era.
On the variance of the fisher information for deep learning
Alexander Soen and Ke Sun. 2021 · 2021
Cited alongside, same era.
A survey on multi-task learning
Yu Zhang and Qiang Yang. 2021 · 2021
Cited alongside, same era.
Mary Phuong and Marcus Hutter. 2022 · 2022
Later among the works it cites.
Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S. Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, and Ludwig Schmidt. 2022 · 2022
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa. 2023 · 2023
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023 · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Xisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, and Pengxiang Cheng. 2023 · 2023
Later among the works it cites.
Model fusion via optimal transport
Sidak Pal Singh and Martin Jaggi. 2023 · 2023
Later among the works it cites.
Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, Hao Peng, Jianxin Li, Jia Wu, Ziwei Liu, Pengtao Xie, Caiming Xiong, Jian Pei, Philip S. Yu, and Lichao Sun. 2023 · 2023
Later among the works it cites.