Fetching the paper…
Reading the bibliography…
We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs.
A new method for mapping optimization problems onto neural networks
C. Peterson and B. Söderberg · 1989
Earlier work this paper cites.
The" softmax" nonlinearity: Derivation using statistical mechanics and useful properties as a multiterminal analog circuit element
I. M. Elfadel and J. L. Wyatt Jr · 1993
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
Publicly available clinical bert embeddings
E. Alsentzer, J. R. Murphy, W. Boag, W.-H. Weng, D. Jin, T. Naumann, and M. McDermott · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
D. Hendrycks and T. Dietterich · 2019
Earlier work this paper cites.
Natural adversarial examples.(2019)
D. Hendrycks, K. Zhao, S. Basart, J. Steinhardt, and D. Song · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Earlier work this paper cites.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, Z. Lipton, and E. P. Xing · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, et al · 2020
Cited alongside, same era.
The many faces of robustness: A critical analysis of out-of-distribution generalization
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al · 2021
Cited alongside, same era.
Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition
Vision-language pre-training with triple contrastive learning
J. Yang, J. Duan, S. Tran, Y. Xu, S. Chanda, L. Chen, B. Zeng, T. Chilimbi, and J. Huang · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Later among the works it cites.
When and why vision-language models behave like bag-of-words models, and what to do about it?
M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou · 2022
Later among the works it cites.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
Regionclip: Region-based language-image pretraining
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S.-C. Huang, L. Shen, M. P. Lungren, and S. Yeung · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi · 2021
Cited alongside, same era.
Valse: A task-independent benchmark for vision and language models centered on linguistic phenomena
L. Parcalabescu, M. Cafagna, L. Muradjan, A. Frank, I. Calixto, and A. Gatt · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Filip: Fine-grained interactive language-image pre-training
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu · 2021
Cited alongside, same era.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Y. Zeng, X. Zhang, and H. Li · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
C. Zhou, C. C. Loy, and B. Dai · 2022
Later among the works it cites.
Evaluating correctness and faithfulness of instruction-following models for question answering
V. Adlakha, P. BehnamGhader, X. H. Lu, N. Meade, and S. Reddy · 2023
Later among the works it cites.
Limitr: Leveraging local information for medical image-text representation
G. Dawidowicz, E. Hirsch, and A. Tal · 2023
Later among the works it cites.
Hiclip: Contrastive language-image pretraining with hierarchy-aware attention
S. Geng, J. Yuan, Y. Tian, Y. Chen, and Y. Zhang · 2023
Later among the works it cites.
Eureka-moments in transformers: Multi-step tasks reveal softmax induced optimization problems
D. T. Hoffmann, S. Schrodi, N. Behrmann, V. Fischer, and T. Brox · 2023
Later among the works it cites.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Later among the works it cites.
Open vocabulary semantic segmentation with patch aligned contrastive learning
J. Mukhoti, T.-Y. Lin, O. Poursaeed, R. Wang, A. Shah, P. H. Torr, and S.-N. Lim · 2023
Later among the works it cites.
R. Paiss, A. Ephrat, O. Tov, S. Zada, I. Mosseri, M. Irani, and T. Dekel · 2023
Later among the works it cites.
Perceptual grouping in contrastive vision-language models
K. Ranasinghe, B. McKinzie, S. Ravi, Y. Yang, A. Toshev, and J. Shlens · 2023
Later among the works it cites.
Dial BeInfo for Faithfulness
E. Razumovskaia, I. Vulić, P. Marković, T. Cichy, Q. Zheng, T.-H. Wen, and P. Budzianowski · 2023
Later among the works it cites.
A study on relu and softmax in transformer
K. Shen, J. Guo, X. Tan, S. Tang, R. Wang, and J. Bian · 2023
Later among the works it cites.
Villa: Fine-grained vision-language representation learning from real-world data
M. Varma, J.-B. Delbrouck, S. Hooper, A. Chaudhari, and C. Langlotz · 2023
Later among the works it cites.
Learning open-vocabulary semantic segmentation models from natural language supervision
J. Xu, J. Hou, Y. Zhang, R. Feng, Y. Wang, Y. Qiao, and W. Xie · 2023
Later among the works it cites.