Branch-train-merge: Embarrassingly parallel training of expert language models
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Later among the works it cites.
Deep neural network fusion via graph matching with applications to model ensemble and federated learning
Liu, C., Lou, C., Wang, R., Xi, A. Y., Shen, L., and Yan, J · 2022
Later among the works it cites.
Merging models with fisher-weighted averaging
Matena, M. S. and Raffel, C · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Exploring mode connectivity for pre-trained language models
Qin, Y., Qian, C., Yi, J., Chen, W., Lin, Y., Han, X., Liu, Z., Sun, M., and Zhou, J · 2022
Later among the works it cites.
Diverse weight averaging for out-of-distribution generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., patrick gallinari, and Cord, M · 2022
Later among the works it cites.
Layerwise linear mode connectivity
Original
Adilova, L., Fischer, A., and Jaggi, M · 2023
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S., Hayase, J., and Srinivasa, S · 2023
Later among the works it cites.
Proving linear mode connectivity of neural networks via optimal transport
Original
Ferbach, D., Goujaud, B., Gidel, G., and Dieuleveut, A · 2023
Later among the works it cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Later among the works it cites.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2023
Later among the works it cites.
Linear connectivity reveals generalization strategies
Juneja, J., Bansal, R., Cho, K., Sedoc, J., and Saphra, N · 2023
Later among the works it cites.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Ortiz-Jimenez, G., Favero, A., and Frossard, P · 2023
Later among the works it cites.
Model ratatouille: Recycling diverse models for out-of-distribution generalization
Rame, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D · 2023
Later among the works it cites.
Zipit! merging models from different tasks without training, 2023
Stoica, G., Bolya, D., Bjorner, J., Hearn, T., and Hoffman, J · 2023
Later among the works it cites.
TIES-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch, 2023
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2023
Later among the works it cites.
Understanding mode connectivity via parameter space symmetry
Zhao, B., Dehmamy, N., Walters, R., and Yu, R · 2023
Later among the works it cites.
Going beyond linear mode connectivity: The layerwise linear feature connectivity
Zhou, Z., Yang, Y., Yang, X., Yan, J., and Hu, W · 2023
Later among the works it cites.
Going beyond neural network feature similarity: The network feature complexity and its interpretation using category theory
Chen, Y., Zhou, Z., and Yan, J · 2024
Closest in time.