Fetching the paper…
Reading the bibliography…
Foundation models serve as the backbone for numerous specialized models developed through fine-tuning.
A shortest augmenting path algorithm for dense and sparse linear assignment problems
Jonker, R. and Volgenant, T · 1988
Earlier work this paper cites.
Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Reading Digits in Natural Images with Unsupervised Feature Learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., Ng, A. Y., et al · 2011
Earlier work this paper cites.
The German Traffic Sign Recognition Benchmark: A multi-class classification competition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2011
Earlier work this paper cites.
Spectral distances of graphs
Jovanović, I. and Stanić, Z · 2012
Earlier work this paper cites.
Describing Textures in the Wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Topology and Geometry of Half-Rectified Network Optimization
Freeman, C. D. and Bruna, J · 2017
Earlier work this paper cites.
Essentially No Barriers in Neural Network Energy Landscape
Draxler, F., Veschgini, K., Salmhofer, M., and Hamprecht, F · 2018
Earlier work this paper cites.
Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
Garipov, T., Izmailov, P., Podoprikhin, D., Vetrov, D. P., and Wilson, A. G · 2018
Earlier work this paper cites.
Averaging Weights Leads to Wider Optima and Better Generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Earlier work this paper cites.
EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2019
Earlier work this paper cites.
GLUE A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Linear Mode Connectivity and the Lottery Ticket Hypothesis
Frankle, J., Dziugaite, G. K., Roy, D., and Carbin, M · 2020
Cited alongside, same era.
Model Fusion via Optimal Transport
Singh, S. P. and Jaggi, M · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., Song, D., Steinhardt, J., and Gilmer, J · 2021
Cited alongside, same era.
OpenCLIP, 2021
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Cited alongside, same era.
REPAIR: REnormalizing Permuted Activations for Interpolation Repair
Jordan, K., Sedghi, H., Saukh, O., Entezari, R., and Neyshabur, B · 2023
Later among the works it cites.
Equivariant Architectures for Learning in Deep Weight Spaces
Navon, A., Shamsian, A., Achituve, I., Fetaya, E., Chechik, G., and Maron, H · 2023
Later among the works it cites.
Re-basin via implicit Sinkhorn differentiation
Peña, F. A. G., Medeiros, H. R., Dubail, T., Aminbeidokhti, M., Granger, E., and Pedersoli, M · 2023
Later among the works it cites.
Model Ratatouille: Recycling Diverse Models for Out-of-Distribution Generalization
Ramé, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D · 2023
Later among the works it cites.
C 2 M 3 : Cycle-Consistent Multi-Model Merging
Crisostomi, D., Fumero, M., Baieri, D., Bernard, F., and Rodola, E · 2024
Later among the works it cites.
Datacomp: In search of the next generation of multimodal datasets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Fusing finetuned models for better pretraining
Choshen, L., Venezian, E., Slonim, N., and Katz, Y · 2022
Cited alongside, same era.
The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks
Entezari, R., Sedghi, H., Saukh, O., and Neyshabur, B · 2022
Cited alongside, same era.
Patching open-vocabulary models by interpolating weights
Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L · 2022
Cited alongside, same era.
Merging Models with Fisher-Weighted Averaging
Matena, M. S. and Raffel, C. A · 2022
Cited alongside, same era.
Diverse Weight Averaging for Out-of-Distribution Generalization
Rame, A., Kirchmeyer, M., Rahier, T., Rakotomamonjy, A., patrick gallinari, and Cord, M · 2022
Cited alongside, same era.
Git Re-Basin: Merging Models modulo Permutation Symmetries
Ainsworth, S., Hayase, J., and Srinivasa, S · 2023
Cited alongside, same era.
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al · 2024
Later among the works it cites.
Transformer Fusion with Optimal Transport
Imfeld, M., Graldi, J., Giordano, M., Hofmann, T., Anagnostidis, S., and Singh, S. P · 2024
Later among the works it cites.
A visual-language foundation model for computational pathology
Lu, M. Y., Chen, B., Williamson, D. F., Chen, R. J., Liang, I., Ding, T., Jaume, G., Odintsov, I., Le, L. P., Gerber, G., et al · 2024
Later among the works it cites.
Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment
Mall, U., Phoo, C. P., Liu, M. K., Vondrick, C., Hariharan, B., and Bala, K · 2024
Later among the works it cites.
Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained Models
Ortiz-Jimenez, G., Favero, A., and Frossard, P · 2024
Later among the works it cites.
ZipIt! Merging Models from Different Tasks without Training
Stoica, G., Bolya, D., Bjorner, J., Ramesh, P., Hearn, T., and Hoffman, J · 2024
Later among the works it cites.
TIES-Merging: Resolving Interference When Merging Models
Yadav, P., Tam, D., Choshen, L., Raffel, C. A., and Bansal, M · 2024
Later among the works it cites.
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
Zhang, F. Z., Albert, P., Rodriguez-Opazo, C., van den Hengel, A., and Abbasnejad, E · 2024
Later among the works it cites.
A Second-Order Perspective on Model Compositionality and Incremental Learning
Porrello, A., Bonicelli, L., Buzzega, P., Millunzi, M., Calderara, S., and Cucchiara, R · 2025
Closest in time.