Fetching the paper…
Reading the bibliography…
Foundation models encode rich representations that can be adapted to downstream tasks by fine-tuning.
Mixout: Effective regularization to finetune large-scale pretrained language models
Lee, C., Cho, K., and Kang, W · 1909
Earlier work this paper cites.
What would elsa do? freezing layers during transformer fine-tuning
Lee, J., Tang, R., and Lin, J · 1911
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Fei-Fei, L., Fergus, R., and Perona, P · 2004
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Torralba, A., Fergus, R., and Freeman, W. T · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Hutter, F., Hoos, H. H., and Leyton-Brown, K · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Bengio, S., Bengio, Y., Cloutier, J., and Gescei, J · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., and Sivic, J · 2014
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
Sharif Razavian, A., Azizpour, H., Sullivan, J., and Carlsson, S · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Tzeng, E., Hoffman, J., Zhang, N., Saenko, K., and Darrell, T · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Efficient and robust automated machine learning
Feurer, M., Klein, A., Eggensperger, K., Springenberg, J., Blum, M., and Hutter, F · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and De Freitas, N · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., and Talwalkar, A · 2017
Earlier work this paper cites.
Learned optimizers that scale and generalize
Wichrowska, O., Maheswaranathan, N., Hoffman, M. W., Colmenarejo, S. G., Denil, M., de Freitas, N., and Sohl-Dickstein, J · 2017
Earlier work this paper cites.
Functional map of the world
Christie, G., Fendley, N., Wilson, J., and Mukherjee, R · 2018
Earlier work this paper cites.
Darts: Differentiable architecture search
Liu, H., Simonyan, K., and Yang, Y · 2018
Earlier work this paper cites.
Do cifar-10 classifiers generalize to cifar-10?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2018
Earlier work this paper cites.
Rotation equivariant cnns for digital pathology
Veeling, B. S., Linmans, J., Winkens, J., Cohen, T., and Welling, M · 2018
Earlier work this paper cites.
Explicit inductive bias for transfer learning with convolutional networks
Xuhong, L., Grandvalet, Y., and Davoine, F · 2018
Earlier work this paper cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
Yu, T., Finn, C., Xie, A., Dasari, S., Zhang, T., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Learning transferable architectures for scalable image recognition
Zoph, B., Vasudevan, V., Shlens, J., and Le, Q. V · 2018
Earlier work this paper cites.
Optuna: A next-generation hyperparameter optimization framework
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M · 2019
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2019
Cited alongside, same era.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B · 2019
Cited alongside, same era.
What is the effect of importance weighting in deep learning?
Byrd, J. and Lipton, Z · 2019
Cited alongside, same era.
Autoaugment: Learning augmentation strategies from data
Cubuk, E. D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V · 2019
Cited alongside, same era.
Spottune: transfer learning through adaptive fine-tuning
Guo, Y., Shi, H., Kumar, A., Grauman, K., Rosing, T., and Feris, R · 2019
Cited alongside, same era.
The evolution of out-of-distribution robustness throughout fine-tuning
Andreassen, A., Bahri, Y., Neyshabur, B., and Roelofs, R · 2021
Later among the works it cites.
Meta learning via learned loss
Bechtle, S., Molchanov, A., Chebotar, Y., Grefenstette, E., Righetti, L., Sukhatme, G., and Meier, F · 2021
Later among the works it cites.
The iwildcam 2021 competition dataset
Beery, S., Agarwal, A., Cole, E., and Birodkar, V · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Later among the works it cites.
Environment inference for invariant learning
Creager, E., Jacobsen, J.-H., and Zemel, R · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hendrycks, D. and Dietterich, T · 2019
Cited alongside, same era.
Using self-supervised learning can improve model robustness and uncertainty
Hendrycks, D., Mazeika, M., Kadavath, S., and Song, D · 2019
Cited alongside, same era.
Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Zhao, T · 2019
Cited alongside, same era.
Improving generalization in meta reinforcement learning using learned objectives
Kirsch, L., van Steenkiste, S., and Schmidhuber, J · 2019
Cited alongside, same era.
Fast autoaugment
Lim, S., Kim, I., Kim, T., Kim, C., and Kim, S · 2019
Cited alongside, same era.
Regularized evolution for image classifier architecture search
Real, E., Aggarwal, A., Huang, Y., and Le, Q. V · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Cited alongside, same era.
Source-free adaptation to measurement shift via bottom-up feature restoration
Eastwood, C., Mason, I., Williams, C. K., and Schölkopf, B · 2021
Later among the works it cites.
Distance-based regularisation of deep networks for fine-tuning
Gouk, H., Hospedales, T., and massimiliano pontil · 2021
Later among the works it cites.
Openclip
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Later among the works it cites.
Test-time adaptable neural networks for robust medical image segmentation
Karani, N., Erdil, E., Chaitanya, K., and Konukoglu, E · 2021
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Marklund, H., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., et al · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., Liang, P., Carmon, Y., and Schmidt, L · 2021
Later among the works it cites.
Partial is better than all: Revisiting fine-tuning strategy for few-shot learning
Shen, Z., Liu, Z., Qin, J., Savvides, M., and Cheng, K.-T · 2021
Later among the works it cites.
A fine-grained analysis on distribution shift
Wiles, O., Gowal, S., Stimberg, F., Alvise-Rebuffi, S., Ktena, I., Cemgil, T., et al · 2021
Later among the works it cites.
Adaptive risk minimization: Learning to adapt to domain shift
Zhang, M., Marklund, H., Dhawan, N., Gupta, A., Levine, S., and Finn, C · 2021
Later among the works it cites.
Agreement-on-the-line: Predicting the performance of neural networks under distribution shift
Baek, C., Jiang, Y., Raghunathan, A., and Kolter, J. Z · 2022
Later among the works it cites.
" this is my unicorn, fluffy": Personalizing frozen vision-language representations
Cohen, N., Gal, R., Meirom, E. A., Chechik, G., and Atzmon, Y · 2022
Later among the works it cites.
Unit-level surprise in neural networks
Eastwood, C., Mason, I., and Williams, C. K · 2022
Later among the works it cites.
Head2toe: Utilizing intermediate representations for better transfer learning
Evci, U., Dumoulin, V., Larochelle, H., and Mozer, M. C · 2022
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G · 2022
Later among the works it cites.
Velo: Training versatile learned optimizers by scaling up
Metz, L., Harrison, J., Freeman, C. D., Merchant, A., Beyer, L., Bradbury, J., Agrawal, N., Poole, B., Mordatch, I., Roberts, A., et al · 2022
Later among the works it cites.
Extending the WILDS benchmark for unsupervised adaptation
Sagawa, S., Koh, P. W., Lee, T., Gao, I., Xie, S. M., Shen, K., Kumar, A., Hu, W., Yasunaga, M., Marklund, H., Beery, S., David, E., Stavness, I., Guo, W., Leskovec, J., Saenko, K., Hashimoto, T., Levine, S., Finn, C., and Liang, P · 2022
Later among the works it cites.
Three things everyone should know about vision transformers
Touvron, H., Cord, M., El-Nouby, A., Verbeek, J., and Jégou, H · 2022
Later among the works it cites.
Improving out-of-distribution robustness via selective augmentation
Yao, H., Wang, Y., Li, S., Zhang, L., Liang, W., Zou, J., and Finn, C · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., et al · 2023
Later among the works it cites.
Out-of-domain robustness via targeted augmentations
Gao, I., Sagawa, S., Koh, P. W., Hashimoto, T., and Liang, P · 2023
Later among the works it cites.
Fine-tuning can cripple your foundation model; preserving features may be the solution
Mukhoti, J., Gal, Y., Torr, P. H., and Dokania, P. K · 2023
Later among the works it cites.
Trainable projected gradient method for robust fine-tuning
Tian, J., He, Z., Dai, X., Ma, C.-Y., Liu, Y.-C., and Kira, Z · 2023
Later among the works it cites.