Fetching the paper…
Reading the bibliography…
Merging various task-specific Transformer-based models trained on different tasks into a single unified model can execute all the tasks concurrently.
Parameter-Efficient Transfer Learning for NLP, June 2019
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 1902
Earlier work this paper cites.
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, March 2019
Hendrycks, D. and Dietterich, T · 1903
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis, July 2020
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 1912
Earlier work this paper cites.
Adaptive Mixtures of Local Experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Pruning Neural Networks at Initialization: Why are We Missing the Mark?, March 2021
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M · 2009
Earlier work this paper cites.
SUN database: Large-scale scene recognition from abbey to zoo
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A · 2010
Earlier work this paper cites.
Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Describing Textures in the Wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network, March 2015
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Understanding Neural Networks Through Deep Visualization, June 2015
Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H · 2015
Earlier work this paper cites.
Convergent Learning: Do different neural networks learn the same representations?, February 2016
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J · 2016
Earlier work this paper cites.
Remote Sensing Image Scene Classification: Benchmark and State of the Art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
Topology and geometry of half-rectified network optimization: 5th International Conference on Learning Representations, ICLR 2017
Daniel Freeman, C. and Bruna, J · 2017
Earlier work this paper cites.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs, October 2018
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D., and Wilson, A. G · 2018
Earlier work this paper cites.
Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2018
Earlier work this paper cites.
Essentially No Barriers in Neural Network Energy Landscape, February 2019
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F. A · 2019
Earlier work this paper cites.
Averaging Weights Leads to Wider Optima and Better Generalization, February 2019
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2019
Earlier work this paper cites.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Optimizing Mode Connectivity via Neuron Alignment
Tatro, N., Chen, P.-Y., Das, P., Melnyk, I., Sattigeri, P., and Lai, R · 2020
Cited alongside, same era.
Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling, November 2021
Benton, G. W., Maddox, W. J., Lotfi, S., and Wilson, A. G · 2021
Cited alongside, same era.
Masked Autoencoders Are Scalable Vision Learners, December 2021
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Cited alongside, same era.
Editing Models with Task Arithmetic, March 2023
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Later among the works it cites.
Dataless Knowledge Fusion by Merging Weights of Language Models, April 2023
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2023
Later among the works it cites.
Merging Decision Transformers: Weight Averaging for Forming Multi-Task Policies, September 2023
Lawson, D. and Qureshi, A. H · 2023
Later among the works it cites.
Deep Model Fusion: A Survey, September 2023
Li, W., Peng, Y., Zhang, M., Ding, L., Hu, H., and Shen, L · 2023
Later among the works it cites.
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts, March 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
BASE Layers: Simplifying Training of Large, Sparse Models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Cited alongside, same era.
Reading Digits in Natural Images with Unsupervised Feature Learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision, February 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models, December 2022
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Cited alongside, same era.
The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks, July 2022
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2022
Cited alongside, same era.
Kaddour, J · 2022
Cited alongside, same era.
Liang, J., He, R., and Tan, T · 2023
Later among the works it cites.
Federated learning on multimodal data: A comprehensive survey
Lin, Y., Gao, Y., Gong, M., Zhang, S., Zhang, Y., and Li, Z · 2023
Later among the works it cites.
Bag of Tricks for Fully Test-Time Adaptation, October 2023
Mounsaveng, S., Chiaroni, F., Boudiaf, M., Pedersoli, M., and Ayed, I. B · 2023
Later among the works it cites.
Improving Heterogeneous Model Reuse by Density Estimation
Tang, A., Luo, Y., Hu, H., He, F., Su, K., Du, B., Chen, Y., and Tao, D · 2023
Later among the works it cites.
Large-scale multi-modal pre-trained models: A comprehensive survey
Wang, X., Chen, G., Qian, G., Gao, P., Wei, X., Wang, Y., Tian, Y., and Gao, W · 2023
Later among the works it cites.
π \pi -tuning: transferring multimodal foundation models with optimal multi-task interpolation
Wu, C., Wang, T., Ge, Y., Lu, Z., Zhou, R., Shan, Y., and Luo, P · 2023
Later among the works it cites.
Resolving Interference When Merging Models, June 2023
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
AdaMerging: Adaptive Model Merging for Multi-Task Learning, October 2023
Yang, E., Wang, Z., Shen, L., Liu, S., Guo, G., Wang, X., and Tao, D · 2023
Later among the works it cites.
Ye, H. and Xu, D · 2023
Later among the works it cites.
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2023
Later among the works it cites.
Learn From Model Beyond Fine-Tuning: A Survey, October 2023
Zheng, H., Shen, L., Tang, A., Luo, Y., Hu, H., Du, B., and Tao, D · 2023
Later among the works it cites.
The life cycle of knowledge in big language models: A survey
Cao, B., Lin, H., Han, X., and Sun, L · 2024
Closest in time.
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Dai, D., Deng, C., Zhao, C., Xu, R. X., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., Xie, Z., Li, Y. K., Huang, P., Luo, F., Ruan, C., Sui, Z., and Liang, W · 2024
Closest in time.
Mixtral of Experts, January 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.