Fetching the paper…
Reading the bibliography…
Vision foundation models have been explored recently to build general-purpose vision systems.
Microsoft COCO: Common Objects in Context
Tsung Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Generation and Comprehension of Unambiguous Object Descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy · 2015
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-To-Phrase Correspondences for Richer Image-To-Sentence Models
Bryan A Plummer, Liwei Wang, · Chris, M Cervantes, Juan C Caicedo, Julia Hockenmaier, Svetlana Lazebnik, B A Plummer, L Wang, J C Caicedo, and J Hockenmaier · 2015
Earlier work this paper cites.
V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi · 2016
Earlier work this paper cites.
Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2017
Earlier work this paper cites.
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2017
Earlier work this paper cites.
the Open Images Dataset V4: Unified Image Classification, Object Detection, and Visual Relationship Detection at Scale
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari · 2018
Earlier work this paper cites.
LVIS: A Dataset for Large Vocabulary Instance Segmentation
Agrim Gupta, Piotr Dollár, and Ross Girshick · 2019
Earlier work this paper cites.
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova Google, and A I Language · 2019
Earlier work this paper cites.
Generalized Intersection over Union: A Metric and a Loss for Bounding Box Regression
Hamid Rezatofighi, Nathan Tsoi, Junyoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Earlier work this paper cites.
Objects365: A Large-Scale, High-Quality Dataset for Object Detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Yu † Gang, Xiangyu Zhang, Jing Li, and Jian Sun · 2019
Earlier work this paper cites.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Drew A. Hudson and Christopher D. Manning · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
End-To-End Object Detection with Transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
Focal Loss for Dense Object Detection
Tsung Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-To-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Emerging Properties in Self-Supervised Vision Transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Earlier work this paper cites.
Learning Transferable Visual Models from Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Scaling up Visual and Vision-Language Representation Learning with Noisy Text Supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Pix2seq: A Language Modeling Framework for Object Detection
Ting Chen, Saurabh Saxena, Lala Li, David J. Fleet, and Geoffrey Hinton · 2021
Cited alongside, same era.
MDETR – Modulated Detection for End-To-End Multi-Modal Understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Cited alongside, same era.
Per-Pixel Classification Is Not All You Need for Semantic Segmentation
Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov · 2021
Cited alongside, same era.
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-To-Sequence Learning Framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Later among the works it cites.
Exploring Plain Vision Transformer Backbones for Object Detection
Yanghao Li, Hanzi Mao, Ross Girshick, and Kaiming He · 2022
Later among the works it cites.
PhraseCut: Language-Based Image Segmentation in the Wild
Chenyun Wu, Zhe Lin, Scott Cohen, Trung Bui, and Subhransu Maji · 2022
Later among the works it cites.
Jeffrey Ouyang, Zhang Jang, Hyun Cho, Xingyi Zhou, and Philipp Krähenbühl · 2022
Later among the works it cites.
Roboflow 100: A Rich, Multi-Domain Object Detection Benchmark
Floriana Ciaglia, Francesco Saverio Zuppichini, Paul Guerrie, Mark McQuade, and Jacob Solawetz · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
an Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Probabilistic Two-Stage Detection
Xingyi Zhou, Vladlen Koltun, Philipp Krähenb, and Krähenb¨ Krähenbühl · 2021
Cited alongside, same era.
Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation
Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, Tsung-Yi Lin, Ekin D. Cubuk, Quoc V. Le, and Barret Zoph · 2021
Cited alongside, same era.
iBOT: Image BERT Pre-Training with Online Tokenizer
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong · 2022
Cited alongside, same era.
Grounded Language-Image Pre-Training
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jianfeng Gao · 2022
Cited alongside, same era.
DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-Training for Open-World Detection
Lewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang, Dan Xu, Wei Zhang, Zhenguo Li, Chunjing Xu, and Hang Xu · 2022
Cited alongside, same era.
Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with Transformers
Zhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, Ping Luo, and Tong Lu · 2022
Cited alongside, same era.
Universal Instance Perception As Object Discovery and Retrieval
Bin Yan, Yi Jiang, Jiannan Wu, Dong Wang, Ping Luo, Zehuan Yuan, and Huchuan Lu · 2023
Closest in time.
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang · 2023
Closest in time.
Generalized Decoding for Pixel, Image, and Language
Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, Nanyun Peng, Lijuan Wang, Yong Jae Lee, and Jianfeng Gao · 2023
Closest in time.
a Simple Framework for Open-Vocabulary Segmentation and Detection
Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianfeng Gao, Jianwei Yang, and Lei Zhang · 2023
Closest in time.
Segment Everything Everywhere All at Once
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Wang, Lijuan Wang, Jianfeng Gao, and Yong Jae Lee · 2023
Closest in time.
Described Object Detection: Liberating Object Detection with Flexible Expressions
Chi Xie, Zhao Zhang, Yixuan Wu, Feng Zhu, Rui Zhao, and Shuang Liang · 2023
Closest in time.
Open-Vocabulary Panoptic Segmentation with Text-To-Image Diffusion Models
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiaolong Wang, and Shalini De Mello · 2023
Closest in time.
Hierarchical Open-Vocabulary Universal Image Segmentation
Xudong Wang, Shufan Li, Konstantinos Kallidromitis, Yusuke Kato, Kazuki Kozuka, and Trevor Darrell · 2023
Closest in time.
Open-Vocabulary Semantic Segmentation with Mask-Adapted CLIP
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu · 2023
Closest in time.
CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-Matching
Xiaoshi Wu, Feng Zhu, Rui Zhao, and Hongsheng Li · 2023
Closest in time.
EVA-02: A Visual Representation for Neon Genesis
Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao · 2023
Closest in time.
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao · 2023
Closest in time.
Kosmos-2: Grounding Multimodal Large Language Models to the World
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei · 2023
Closest in time.