Fetching the paper…
Reading the bibliography…
Transfer learning with a small amount of target data is an effective and common approach to adapting a pre-trained model to distribution shifts.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D. (2019) · 1907
Earlier work this paper cites.
Semenova, L., Rudin, C., and Parr, R. (2019) · 1908
Earlier work this paper cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al. (2019) · 1910
Earlier work this paper cites.
Feature selection for classification
Dash, M. and Liu, H. (1997) · 1997
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C., and Bialek, W. (2000) · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S. (2002) · 2002
Earlier work this paper cites.
On the relation between discriminant analysis and mutual information for supervised linear feature extraction
Petridis, S. and Perantonis, S. J. (2004) · 2004
Earlier work this paper cites.
Tent: Fully test-time adaptation by entropy minimization
Wang, D., Shelhamer, E., Liu, S., Olshausen, B., and Darrell, T. (2020) · 2006
Earlier work this paper cites.
Nonlinear dimensionality reduction
Lee, J. A., Verleysen, M., et al. (2007) · 2007
Earlier work this paper cites.
Computational methods of feature selection
Liu, H. and Motoda, H. (2007) · 2007
Earlier work this paper cites.
On the complexity of linear prediction: Risk bounds, margin bounds, and regularization
Kakade, S. M., Sridharan, K., and Tewari, A. (2008) · 2008
Earlier work this paper cites.
Breeds: Benchmarks for subpopulation shift
Santurkar, S., Tsipras, D., and Madry, A. (2020) · 2008
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. (2020) · 2010
Earlier work this paper cites.
A survey on feature selection methods
Chandrashekar, G. and Sahin, F. (2014) · 2014
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., and Sivic, J. (2014) · 2014
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
Sharif Razavian, A., Azizpour, H., Sullivan, J., and Carlsson, S. (2014) · 2014
Earlier work this paper cites.
A survey of dimensionality reduction techniques
Sorzano, C. O. S., Vargas, J., and Montano, A. P. (2014) · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Tzeng, E., Hoffman, J., Zhang, N., Saenko, K., and Darrell, T. (2014) · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H. (2014) · 2014
Earlier work this paper cites.
Linear dimensionality reduction: Survey, insights, and generalizations
Cunningham, J. P. and Ghahramani, Z. (2015) · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Liu, Z., Luo, P., Wang, X., and Tang, X. (2015) · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alemi, A. A., Fischer, I., Dillon, J. V., and Murphy, K. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., et al. (2017) · 2017
Cited alongside, same era.
Feature selection: A data perspective
Li, J., Cheng, K., Wang, S., Morstatter, F., Trevino, R. P., Tang, J., and Liu, H. (2017) · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Cited alongside, same era.
Wilds: A benchmark of in-the-wild distribution shifts
Koh, P. W., Sagawa, S., Xie, S. M., Zhang, M., Balsubramani, A., Hu, W., Yasunaga, M., Phillips, R. L., Gao, I., Lee, T., et al. (2021) · 2021
Later among the works it cites.
Just train twice: Improving group robustness without training group information
Liu, E. Z., Haghgoo, B., Chen, A. S., Raghunathan, A., Koh, P. W., Sagawa, S., Liang, P., and Finn, C. (2021) · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, S.-O., Bengio, Y., Courville, A., Precup, D., and Lajoie, G. (2021) · 2021
Later among the works it cites.
Teney, D., Abbasnejad, E., Lucey, S., and Hengel, A. v. d. (2021) · 2021
Later among the works it cites.
Adaptive risk minimization: Learning to adapt to domain shift
Zhang, M., Marklund, H., Dhawan, N., Gupta, A., Levine, S., and Finn, C. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
From detection of individual metastases to classification of lymph node status at the patient level: the camelyon17 challenge
Bandi, P., Geessink, O., Manson, Q., Van Dijk, M., Balkenhol, M., Hermsen, M., Bejnordi, B. E., Lee, B., Paeng, K., Zhong, A., et al. (2018) · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Gunasekar, S., Lee, J. D., Soudry, D., and Srebro, N. (2018) · 2018
Cited alongside, same era.
Do better imagenet models transfer better? arxiv 2018
Kornblith, S., Shlens, J., and Le, Q. (2018) · 2018
Cited alongside, same era.
All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously
Fisher, A., Rudin, C., and Dominici, F. (2019) · 2019
Cited alongside, same era.
On the downstream performance of compressed word embeddings
May, A., Zhang, J., Dao, T., and Ré, C. (2019) · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint
Wainwright, M. J. (2019) · 2019
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. (2020) · 2020
Cited alongside, same era.
Test-time training with masked autoencoders
Gandelsman, Y., Sun, Y., Chen, X., and Efros, A. A. (2022) · 2022
Later among the works it cites.
On feature learning in the presence of spurious correlations
Izmailov, P., Kirichenko, P., Gruver, N., and Wilson, A. G. (2022) · 2022
Later among the works it cites.
Last layer re-training is sufficient for robustness to spurious correlations
Kirichenko, P., Izmailov, P., and Wilson, A. G. (2022) · 2022
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P. (2022) · 2022
Later among the works it cites.
A whac-a-mole dilemma: Shortcuts come in multiples where mitigating one amplifies others
Li, Z., Evtimov, I., Gordo, A., Hazirbas, C., Hassner, T., Ferrer, C. C., Xu, C., and Ibrahim, M. (2022) · 2022
Later among the works it cites.
Lubana, E. S., Bigelow, E. J., Dick, R. P., Krueger, D., and Tanaka, H. (2022) · 2022
Later among the works it cites.
You only need a good embeddings extractor to fix spurious correlations
Mehta, R., Albiero, V., Chen, L., Evtimov, I., Glaser, T., Li, Z., and Hassner, T. (2022) · 2022
Later among the works it cites.
Agree to disagree: Diversity through disagreement for better transferability
Pagliardini, M., Jaggi, M., Fleuret, F., and Karimireddy, S. P. (2022) · 2022
Later among the works it cites.
Rosenfeld, E., Ravikumar, P., and Risteski, A. (2022) · 2022
Later among the works it cites.
Evading the simplicity bias: Training a diverse set of models discovers solutions with superior ood generalization
Teney, D., Abbasnejad, E., Lucey, S., and van den Hengel, A. (2022) · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L. (2022) · 2022
Later among the works it cites.
Controlling directions orthogonal to a classifier
Xu, Y., He, H., Shen, T., and Jaakkola, T. (2022) · 2022
Later among the works it cites.
Contrastive adapters for foundation model group robustness
Zhang, M. and Ré, C. (2022) · 2022
Later among the works it cites.
Simplicity bias in 1-hidden layer neural networks
Morwani, D., Batra, J., Jain, P., and Netrapalli, P. (2023) · 2023
Closest in time.
Domain-adversarial training of neural networks
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., and Lempitsky, V. (2016) · 2030
Closest in time.