Fetching the paper…
Reading the bibliography…
Transfer learning plays a key role in advancing machine learning models, yet conventional supervised pretraining often undermines feature transferability by prioritizing features that minimize the pretraining loss.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Efficient backprop
LeCun, Y., Bottou, L., Orr, G. B., and Müller, K.-R · 2002
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q · 2010
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T · 2011
Earlier work this paper cites.
Deep learning of representations for unsupervised and transfer learning
Bengio, Y · 2012
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C · 2012
Earlier work this paper cites.
3d object representations for fine-grained categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
Maji, S., Kannala, J., Rahtu, E., Blaschko, M., and Vedaldi, A · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A · 2014
Earlier work this paper cites.
Reducing overfitting in deep networks by decorrelating representations
Cogswell, M., Ahmed, F., Girshick, R., Zitnick, L., and Batra, D · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
A survey of transfer learning
Weiss, K. R., Khoshgoftaar, T. M., and Wang, D · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization. corr abs/1611.03530 (2016)
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Cited alongside, same era.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Cited alongside, same era.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Later among the works it cites.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Bardes, A., Ponce, J., and LeCun, Y · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Later among the works it cites.
Why do better loss functions lead to less transferable features?
Kornblith, S., Chen, T., Lee, H., and Norouzi, M · 2021
Later among the works it cites.
Gradient starvation: A learning proclivity in neural networks
Pezeshki, M., Kaba, O., Bengio, Y., Courville, A. C., Precup, D., and Lajoie, G · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Belghazi, M. I., Baratin, A., Rajeswar, S., Ozair, S., Bengio, Y., Courville, A., and Hjelm, R. D · 2018
Cited alongside, same era.
Optimal whitening and decorrelation
Kessy, A., Lewin, A., and Strimmer, K · 2018
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Cited alongside, same era.
Representation compression and generalization in deep neural networks, 2018
Shwartz-Ziv, R., Painsky, A., and Tishby, N · 2018
Cited alongside, same era.
The inaturalist species classification and detection dataset
Van Horn, G., Mac Aodha, O., Song, Y., Cui, Y., Sun, C., Shepard, A., Adam, H., Perona, P., and Belongie, S · 2018
Cited alongside, same era.
Regularizing deep neural networks by enhancing diversity in feature extraction
Ayinde, B. O., Inanc, T., and Zurada, J. M · 2019
Cited alongside, same era.
Robustness (python library), 2019
Engstrom, L., Ilyas, A., Santurkar, S., and Tsipras, D · 2019
Cited alongside, same era.
Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Geiping, J., Goldblum, M., Somepalli, G., Shwartz-Ziv, R., Goldstein, T., and Wilson, A. G · 2022
Later among the works it cites.
Uniformerv2: Spatiotemporal learning by arming image vits with video uniformer, 2022
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Wang, L., and Qiao, Y · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Information flow in deep neural networks
Shwartz-Ziv, R · 2022
Later among the works it cites.
What do we maximize in self-supervised learning?
Shwartz-Ziv, R., Balestriero, R., and LeCun, Y · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Later among the works it cites.
Reverse engineering self-supervised learning
Ben-Shaul, I., Shwartz-Ziv, R., Galanti, T., Dekel, S., and LeCun, Y · 2023
Closest in time.
Wld-reg: A data-dependent within-layer diversity regularizer
Laakom, F., Raitoharju, J., Iosifidis, A., and Gabbouj, M · 2023
Closest in time.
To compress or not to compress–self-supervised learning and information theory: A review
Shwartz-Ziv, R. and LeCun, Y · 2023
Closest in time.
An information-theoretic perspective on variance-invariance-covariance regularization
Shwartz-Ziv, R., Balestriero, R., Kawaguchi, K., Rudner, T. G., and LeCun, Y · 2023
Closest in time.
Videomae v2: Scaling video masked autoencoders with dual masking
Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., and Qiao, Y · 2023
Closest in time.