Fetching the paper…
Reading the bibliography…
While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks.
Distinctive Image Features from Scale-Invariant Keypoints
D. Lowe · 2004
Earlier work this paper cites.
Histograms of Oriented Gradients for Human Detection
N. Dalal and B. Triggs · 2005
Earlier work this paper cites.
The Pascal Visual Object Classes (VOC) Challenge
M. Everingham, L. Van Gool, C. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Imagenet Classification with Deep Convolutional Neural Networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Earlier work this paper cites.
Indoor Segmentation and Support Inference from RGBD Images
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
J. Krause, M. Stark, J. Deng, and L. Fei-Fei · 2013
Earlier work this paper cites.
From Image Descriptions to Visual Denotations: New Similarity Metrics for Semantic Inference over Event Descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Networ
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Learning Visual Features from Large Weakly Supervised Data
A. Joulin, L. van der Maaten, A. Jabri, and N. Vasilache · 2016
Earlier work this paper cites.
DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations
Z. Liu, P. Luo, S. Qiu, X. Wang, and X. Tang · 2016
Earlier work this paper cites.
Deep Metric Learning via Lifted Structured Feature Embedding
H. Song, Y. Xiang, S. Jegelka, and S. Savarese · 2016
Earlier work this paper cites.
Attention is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Scene Parsing Through ADE20k Dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Earlier work this paper cites.
Exploring the Limits of Weakly Supervised Pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Earlier work this paper cites.
Cross-domain Self-supervised Multi-task Feature Learning Using Synthetic Imagery
Z. Ren and Y. J. Lee · 2018
Earlier work this paper cites.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
N. Shazeer and M. Stern · 2018
Earlier work this paper cites.
Representation Learning with Contrastive Predictive Coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
The iNaturalist Species Classification and Detection Dataset
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie · 2018
Earlier work this paper cites.
A Simple Framework for Contrastive Learning of Visual Representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Earlier work this paper cites.
Bootstrap Your Own Latent - A New Approach to Self-Supervised Learning
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, and M. Gheshlaghi Azar · 2020
Earlier work this paper cites.
Momentum Contrast for Unsupervised Visual Representation Learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Earlier work this paper cites.
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
B. Mildenhall, P. Srinivasan, M. Tancik, J. Barron, R. Ramamoorthi, and R. Ng · 2020
Earlier work this paper cites.
RP2K: A Large-Scale Retail Product Dataset forFine-Grained Image Classification
J. Peng, C. Xiao, and Y. Li · 2020
Earlier work this paper cites.
GLU Variants Improve Transformer
N. Shazeer · 2020
Earlier work this paper cites.
Mapillary Street-Level Sequences: A Dataset for Lifelong Place Recognition
F. Warburg, S. Hauberg, M. López-Antequera, P. Gargallo, Y. Kuang, and J. Civera · 2020
Earlier work this paper cites.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
G. Wenzek, M. Lachaux, A. Conneau, V. Chaudhary, F. Guzman, A. Joulin, and E. Grave · 2020
Earlier work this paper cites.
Google Landmarks Dataset v2 - A Large-Scale Benchmark for Instance-Level Recognition and Retrieval
T. Weyand, A. Araujo, B. Cao, and J. Sim · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers
M. Caron, H. Touvron, I. Misra, H. Jegou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
SLIP: Self-supervision meets Language-Image Pre-training
N. Mu, A. Kirillov, D. Wagner, and S. Xie · 2021
Cited alongside, same era.
Learning Transferable Visual Models from Natural Language Supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Scaling Language-Image Pre-training via Masking
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He · 2023
Later among the works it cites.
Large Scale Visual Food Recognition
W. Min, Z. Wang, Y. Liu, M. Luo, L. Kang, X. Wei, X. Wei, and S. Jiang · 2023
Later among the works it cites.
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning
J. Mukhoti, T.-Y. Lin, O. Poursaeed, R. Wang, A. Shah, P. Torr, and S.-N. Lim · 2023
Later among the works it cites.
Self-supervised Video Pretraining Yields Human-aligned Visual Representations
N. Parthasarathy, S. M. Eslami, J. Carreira, and O. Henaff · 2023
Later among the works it cites.
Fake It Till You Make It: Learning Transferable Representations from Synthetic ImageNet Clones
M. B. Sariyildiz, K. Alahari, D. Larlus, and Y. Kalantidis · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vision Transformers for Dense Prediction
R. Ranftl, A. Bochkovskiy, and V. Koltun · 2021
Cited alongside, same era.
The Met Dataset: Instance-level Recognition for Artworks
N.-A. Ypsilantis, N. Garcia, G. Han, S. Ibrahimi, N. Van Noord, and G. Tolias · 2021
Cited alongside, same era.
Florence: A New Foundation Model for Computer Vision
L. Yuan, D. Chen, Y. Chen, N. Codella, X. Dai, J. Gao, H. Hu, X. Huang, B. Li, C. Li, C. Liu, M. Liu, Z. Liu, Y. Lu, Y. Shi, L. Wang, J. Wang, B. Xiao, Z. Xiao, J. Yang, M. Zeng, L. Zhou, and P. Zhang · 2021
Cited alongside, same era.
Masked Autoencoders are Scalable Vision Learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollar, and R. Girshick · 2022
Cited alongside, same era.
Simple Open-Vocabulary Object Detection with Vision Transformers
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, X. Wang, X. Zhai, T. Kipf, and N. Houlsby · 2022
Cited alongside, same era.
DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
Y. Rao, W. Zhao, G. Chen, Y. Tang, Z. Zhu, G. Huang, J. Zhou, and J. Lu · 2022
Cited alongside, same era.
LAION-5B: An Open Large-Scale Dataset for Training Next Generation Image-Text Models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Cited alongside, same era.
Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao · 2023
Later among the works it cites.
StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners
Y. Tian, L. Fan, P. Isola, H. Chang, and D. Krishnan · 2023
Later among the works it cites.
Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations
N.-A. Ypsilantis, K. Chen, B. Cao, M. Lipovský, P. Dogan-Schönberger, G. Makosa, B. Bluntschli, M. Seyedhosseini, O. Chum, and A. Araujo · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-Training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
PaliGemma: A versatile 3B VLM for transfer
L. Beyer, A. Steiner, A. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bošnjak, X. Chen, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai · 2024
Closest in time.
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, B. Li, P. Luo, T. Lu, Y. Qiao, and J. Dai · 2024
Closest in time.
Vision Transformers Need Registers
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski · 2024
Closest in time.
Probing the 3D Awareness of Visual Foundation Models
M. El Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani · 2024
Closest in time.
EVA-02: A Visual Representation for Neon Genesis
Y. Fang, Q. Sun, X. Wang, T. Huang, X. Wang, and Y. Cao · 2024
Closest in time.
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
H. Hammoud, H. Itani, F. Pizzati, P. Torr, A. Bibi, and B. Ghanem · 2024
Closest in time.
LRM: Large Reconstruction Model for Single Image to 3D
Y. Hong, K. Zhang, J. Gu, S. Bi, Y. Zhou, D. Liu, F. Liu, K. Sunkavalli, T. Bui, and H. Tan · 2024
Closest in time.
VeCLIP: Improving CLIP Training via Visual-enriched Captions
Z. Lai, H. Zhang, B. Zhang, W. Wu, H. Bai, A. Timofeev, X. Du, Z. Gan, J. Shan, C. Chuah, Y. Yang, and M. Cao · 2024
Closest in time.
Binsformer: Revisiting Adaptive Bins for Monocular Depth Estimation
Z. Li, X. Wang, X. Liu, and J. Jiang · 2024
Closest in time.
You Don’t Need Data-Augmentation in Self-Supervised Learning
T. Moutakanni, M. Oquab, M. Szafraniec, M. Vakalopoulou, and P. Bojanowski · 2024
Closest in time.
SILC: Improving Vision Language Pretraining with Self-Distillation
M. F. Naeem, Y. Xian, X. Zhai, L. Hoyer, L. Van Gool, and F. Tombari · 2024
Closest in time.
DOCCI: Descriptions of Connected and Contrasting Images
Y. Onoe, S. Rane, Z. Berger, Y. Bitton, J. Cho, R. Garg, A. Ku, Z. Parekh, J. Pont-Tuset, G. Tanzer, S. Wang, and J. Baldridge · 2024
Closest in time.
DINOv2: Learning Robust Visual Features without Supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, R. Howes, P. Huang, H. Xu, V. Sharma, S. Li, W. Galuba, M. Rabbat, M. Assran, N. Ballas, G. Synnaeve, I. Misra, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Closest in time.
Learning Vision from Models Rivals Learning Vision from Data
Y. Tian, L. Fan, K. Chen, D. Katabi, D. Krishnan, and P. Isola · 2024
Closest in time.
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie · 2024
Closest in time.
PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction
P. Wang, H. Tan, S. Bi, Y. Xu, F. Luan, K. Sunkavalli, W. Wang, Z. Xu, and K. Zhang · 2024
Closest in time.
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
S. Wu, W. Zhang, L. Xu, S. Jin, X. Li, W. Liu, and C. C. Loy · 2024
Closest in time.
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
M. Wysoczanska, O. Simeoni, M. Ramamonjisoa, A. Bursuc, T. Trzcinski, and P. Perez · 2024
Closest in time.
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao · 2024
Closest in time.
CapsFusion: Rethinking Image-Text Data at Scale
Q. Yu, Q. Sun, X. Zhang, Y. Cui, F. Zhang, Y. Cao, X. Wang, and J. Liu · 2024
Closest in time.