Fetching the paper…
Reading the bibliography…
Parameter-efficient tuning has become a trend in transferring large-scale foundation models to downstream applications.
Learning methods for generic object recognition with invariance to pose and lighting
Y. LeCun, F. J. Huang, and L. Bottou · 2004
Earlier work this paper cites.
One-shot learning of object categories
F.-F. Li, R. Fergus, and P. Perona · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
SUN database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
Novel dataset for fine-grained image categorization: Stanford dogs
A. Khosla, N. Jayadevaprakash, B. Yao, and F.-F. Li · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Vision meets robotics: The KITTI dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Fine-grained visual classification of aircraft
S. Maji, J. Kannala, E. Rahtu, M. Blaschko, and A. Vedaldi · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Food-101 – mining discriminative components with random forests
L. Bossard, M. Guillaumin, and L. Van Gool · 2014
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, , and A. Vedaldi · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Kaggle diabetic retinopathy detection competition report., 2015
B. Graham · 2015
Earlier work this paper cites.
U-Net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Earlier work this paper cites.
Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection
G. Van Horn, S. Branson, R. Farrell, S. Haber, J. Barry, P. Ipeirotis, P. Perona, and S. Belongie · 2015
Earlier work this paper cites.
C. Beattie, J. Z. Leibo, D. Teplyashin, T. Ward, M. Wainwright, H. Küttler, A. Lefrancq, S. Green, V. Valdés, A. Sadik, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
G. Cheng, J. Han, and X. Lu · 2017
Earlier work this paper cites.
Fine-grained car detection for visual census estimation
T. Gebru, J. Krause, Y. Wang, D. Chen, J. Deng, and L. Fei-Fei · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Earlier work this paper cites.
dSprites: Disentanglement testing sprites dataset, 2017
L. Matthey, I. Higgins, D. Hassabis, and A. Lerchner · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
BERT: Pre-training of deep bidirectional Transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Rotation equivariant cnns for digital pathology
B. S. Veeling, J. Linmans, J. Winkens, T. Cohen, and M. Welling · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. Bowman · 2018
Cited alongside, same era.
Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
P. Helber, B. Bischke, A. Dengel, and D. Borth · 2019
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Pyramid vision Transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Later among the works it cites.
BitFit: Simple parameter-efficient fine-tuning for Transformer-based masked language-models
E. B. Zaken, S. Ravfogel, and Y. Goldberg · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with Transformers
S. Zheng, J. Lu, H. Zhao, X. Zhu, Z. Luo, Y. Wang, Y. Fu, J. Feng, T. Xiang, P. H. Torr, and L. Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
VisualBERT: A simple and performant baseline for vision and language
L. H. Li, M. Yatskar, D. Yin, C.-J. Hsieh, and K.-W. Chang · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Cited alongside, same era.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. C. Lipton, and E. P. Xing · 2019
Cited alongside, same era.
A large-scale study of representation learning with the visual task adaptation benchmark
X. Zhai, J. Puigcerver, A. Kolesnikov, P. Ruyssen, C. Riquelme, M. Lucic, J. Djolonga, A. S. Pinto, M. Neumann, A. Dosovitskiy, et al · 2019
Cited alongside, same era.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Later among the works it cites.
AdaptFormer: Adapting vision Transformers for scalable visual recognition
S. Chen, C. Ge, Z. Tong, J. Wang, Y. Song, J. Wang, and P. Luo · 2022
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao · 2022
Later among the works it cites.
Visual prompt tuning
M. Jia, L. Tang, B.-C. Chen, C. Cardie, S. Belongie, B. Hariharan, and S.-N. Lim · 2022
Later among the works it cites.
Convolutional bypasses are better vision Transformer adapters
S. Jie and Z.-H. Deng · 2022
Later among the works it cites.
UniFormer: Unified Transformer for efficient spatiotemporal representation learning
K. Li, Y. Wang, P. Gao, G. Song, Y. Liu, H. Li, and Y. Qiao · 2022
Later among the works it cites.
Scaling & shifting your features: A new baseline for efficient model tuning
D. Lian, D. Zhou, J. Feng, and X. Wang · 2022
Later among the works it cites.
Trackformer: Multi-object tracking with transformers
T. Meinhardt, A. Kirillov, L. Leal-Taixe, and C. Feichtenhofer · 2022
Later among the works it cites.
Scalable diffusion models with Transformers
W. Peebles and S. Xie · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with CLIP latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman · 2022
Later among the works it cites.
ViDT: An efficient and effective fully Transformer-based object detector
H. Song, D. Sun, S. Chun, V. Jampani, D. Han, B. Heo, W. Kim, and M.-H. Yang · 2022
Later among the works it cites.
LST: Ladder side-tuning for parameter and memory efficient transfer learning
Y.-L. Sung, J. Cho, and M. Bansal · 2022
Later among the works it cites.
Scaling Vision Transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
Y. Zhang, K. Zhou, and Z. Liu · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
K. Zhou, J. Yang, C. C. Loy, and Z. Liu · 2022
Later among the works it cites.
Scaling Vision Transformers to 22 billion parameters
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, et al · 2023
Closest in time.