Data determines distributional robustness in contrastive language image pre-training (CLIP)
Fang, A., Ilharco, G., Wortsman, M., Wan, Y., Shankar, V., Dave, A., and Schmidt, L · 2022
Closest in time.
No one representation to rule them all: Overlapping features of training methods
Gontijo-Lopes, R., Dauphin, Y., and Cubuk, E. D · 2022
Closest in time.
Patching open-vocabulary models by interpolating weights
Ilharco, G., Wortsman, M., Gadre, S. Y., Song, S., Hajishirzi, H., Kornblith, S., Farhadi, A., and Schmidt, L · 2022
Closest in time.
Combining diverse feature priors
Jain, S., Tsipras, D., and Madry, A · 2022
Closest in time.
Stop wasting my time! saving days of imagenet and BERT training with latest weight averaging
Kaddour, J · 2022
Closest in time.
Plex: Towards reliability using pretrained large model extensions
Kirsch, A. C., Lakshminarayanan, B., Hu, C. H., Sculley, D., Phan, D., Tran, D., Snoek, J. R., Liu, J., Ren, J. J., van Amersfoort, J., Han, K., Buchanan, K., Murphy, K. P., Collier, M. P., Dusenberry, M. W., Band, N., Thain, N., Jenatton, R., Rudner, T. G. J., Gal, Y., Nado, Z., Mariet, Z., Wang, Z., and Ghahramani, Z · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R. M., Ma, T., and Liang, P · 2022
Closest in time.
BERT WEAVER: Using WEight AVERaging to enable lifelong learning for transformer-based models
Langnickel, L., Schulz, A., Hammer, B., and Fluck, J · 2022
Closest in time.
Measuring and signing fairness as performance under multiple stakeholder distributions
Lopez-Paz, D., Bouchacourt, D., Sagun, L., and Usunier, N · 2022
Closest in time.
Model soups improve performance of dermoscopic skin cancer classifiers
Maron, R. C., Hekler, A., Haggenmüller, S., von Kalle, C., Utikal, J. S., Müller, V., Gaiser, M., Meier, F., Hobelsberger, S., Gellrich, F. F., et al · 2022
Closest in time.
Merging models with Fisher-weighted averaging
Matena, M. and Raffel, C · 2022
Closest in time.
Diverse ImageNet models transfer better
Nayman, N., Golbert, A., Noy, A., Ping, T., and Zelnik-Manor, L · 2022
Closest in time.
Quality not quantity: On the interaction between dataset design and robustness of CLIP
Nguyen, T., Ilharco, G., Wortsman, M., Oh, S., and Schmidt, L · 2022
Closest in time.
Exploring mode connectivity for pre-trained language models
Qin, Y., Qian, C., Yi, J., Chen, W., Lin, Y., Han, X., Liu, Z., Sun, M., and Zhou, J · 2022
Closest in time.
Momentum-based weight interpolation of strong zero-shot models for continual learning
Stojanovski, Z., Roth, K., and Akata, Z · 2022
Closest in time.
ID and OOD performance are sometimes inversely correlated on real-world datasets
Teney, D., Lin, Y., Oh, S. J., and Abbasnejad, E · 2022
Closest in time.
Assaying out-of-distribution generalization in transfer learning
Wenzel, F., Dittadi, A., Gehler, P. V., Simon-Gabriel, C.-J., Horn, M., Zietlow, D., Kernert, D., Russell, C., Brox, T., Schiele, B., Schölkopf, B., and Locatello, F · 2022
Closest in time.
Ood-bench: Benchmarking and understanding out-of-distribution generalization datasets and algorithms
Ye, N., Li, K., Hong, L., Bai, H., Chen, Y., Zhou, F., and Li, Z · 2022
Closest in time.
Learning useful representations for shifting tasks and distributions
Zhang, J. and Bottou, L · 2022
Closest in time.
Rich feature construction for the optimization-generalization dilemma
Zhang, J., Lopez-Paz, D., and Bottou, L · 2022
Closest in time.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2023
Closest in time.
Does progress on ImageNet transfer to real-world datasets?
Fang, A., Kornblith, S., and Schmidt, L · 2023
Closest in time.
Editing models with task arithmetic
Ilharco, G., Tulio Ribeiro, M., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Closest in time.
Dataless knowledge fusion by merging weights of language models
Jin, X., Ren, X., Preotiuc-Pietro, D., and Cheng, P · 2023
Closest in time.
Repair: Renormalizing permuted activations for interpolation repair
Jordan, K., Sedghi, H., Saukh, O., Entezari, R., and Neyshabur, B · 2023
Closest in time.
Linear connectivity reveals generalization strategies
Juneja, J., Bansal, R., Cho, K., Sedoc, J., and Saphra, N · 2023
Closest in time.
Building machine learning models like open source software
Raffel, C · 2023
Closest in time.
Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Original
Ramé, A., Couairon, G., Shukor, M., Dancette, C., Gaya, J.-B., Soulier, L., and Cord, M · 2023
Closest in time.
Unified model for image, video, audio and language tasks
Original
Shukor, M., Dancette, C., Ramé, A., and Cord, M · 2023
Closest in time.
lo-fi: distributed fine-tuning without communication
Wortsman, M., Gururangan, S., Li, S., Farhadi, A., Schmidt, L., Rabbat, M., and Morcos, A. S · 2023
Closest in time.