Fetching the paper…
Reading the bibliography…
From image-text pairs, large-scale vision-language models (VLMs) learn to implicitly associate image regions with words, which prove effective for tasks like visual question answering.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Philipp Krähenbühl and Vladlen Koltun · 2011
Earlier work this paper cites.
Unsupervised joint object discovery and segmentation in internet images
Michael Rubinstein, Armand Joulin, Johannes Kopf, and Ce Liu · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille · 2014
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2014
Earlier work this paper cites.
Unsupervised object discovery and localization in the wild: Part-based matching with bottom-up region proposals
Minsu Cho, Suha Kwak, Cordelia Schmid, and Jean Ponce · 2015
Earlier work this paper cites.
Self-taught object localization with deep networks
Loris Bazzani, Alessandra Bergamo, Dragomir Anguelov, and Lorenzo Torresani · 2016
Earlier work this paper cites.
Weakly supervised object localization with multi-fold multiple instance learning
Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid · 2016
Earlier work this paper cites.
Seed, expand and constrain: Three principles for weakly-supervised image segmentation
Alexander Kolesnikov and Christoph H Lampert · 2016
Earlier work this paper cites.
Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization
Krishna Kumar Singh and Yong Jae Lee · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Object region mining with adversarial erasing: A simple classification to semantic segmentation approach
Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari · 2018
Earlier work this paper cites.
A review of semantic segmentation using deep neural networks
Yanming Guo, Yu Liu, Theodoros Georgiou, and Michael S Lew · 2018
Earlier work this paper cites.
Self-erasing network for integral object attention
Qibin Hou, PengTao Jiang, Yunchao Wei, and Ming-Ming Cheng · 2018
Earlier work this paper cites.
On regularized losses for weakly-supervised cnn segmentation
Meng Tang, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, and Yuri Boykov · 2018
Earlier work this paper cites.
Adversarial complementary learning for weakly supervised object localization
Xiaolin Zhang, Yunchao Wei, Jiashi Feng, Yi Yang, and Thomas S Huang · 2018
Earlier work this paper cites.
Zero-shot semantic segmentation
Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Pérez · 2019
Earlier work this paper cites.
Unsupervised image matching and object discovery as optimization
Huy V Vo, Francis Bach, Minsu Cho, Kai Han, Yann LeCun, Patrick Pérez, and Jean Ponce · 2019
Earlier work this paper cites.
Unsupervised object discovery and co-localization by deep descriptor transformation
Xiu-Shen Wei, Chen-Lin Zhang, Jianxin Wu, Chunhua Shen, and Zhi-Hua Zhou · 2019
Earlier work this paper cites.
Semantic projection network for zero-and few-label semantic segmentation
Yongqin Xian, Subhabrata Choudhury, Yang He, Bernt Schiele, and Zeynep Akata · 2019
Cited alongside, same era.
Single-stage semantic segmentation from image labels
Nikita Araslanov and Stefan Roth · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Context-aware feature generation for zero-shot semantic segmentation
Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang · 2020
Cited alongside, same era.
A brief survey on semantic segmentation with deep learning
Shijie Hao, Yuan Zhou, and Yanrong Guo · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao · 2020
Open-world semantic segmentation via contrasting and clustering vision-language embedding
Quande Liu, Youpeng Wen, Jianhua Han, Chunjing Xu, Hang Xu, and Xiaodan Liang · 2022
Later among the works it cites.
Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers
Lixiang Ru, Yibing Zhan, Baosheng Yu, and Bo Du · 2022
Later among the works it cites.
Reco: Retrieve and co-segment for zero-shot transfer
Gyungin Shin, Weidi Xie, and Samuel Albanie · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
Plug-and-play vqa: Zero-shot vqa by conjoining large pretrained models with zero training
Anthony Meng Huat Tiong, Junnan Li, Boyang Li, Silvio Savarese, and Steven C.H. Hoi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient-Free-Optimizers: Simple and reliable optimization with local, global, population-based and sequential techniques in numerical search spaces
Simon Blanke · 2020
Cited alongside, same era.
Reliability does matter: An end-to-end weakly supervised semantic segmentation approach
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang · 2020
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Universal weakly supervised segmentation by pixel-to-segment contrastive learning
Tsung-Wei Ke, Jyh-Jing Hwang, and Stella X Yu · 2021
Cited alongside, same era.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Cited alongside, same era.
A closer look at self-training for zero-label semantic segmentation
Giuseppe Pastore, Fabio Cermelli, Yongqin Xian, Massimiliano Mancini, Zeynep Akata, and Barbara Caputo · 2021
Cited alongside, same era.
CoCa: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Open-vocabulary semantic segmentation using test-time distillation
Nir Zabari and Yedid Hoshen · 2022
Later among the works it cites.
Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs
Junbum Cha, Jonghwan Mun, and Byungseok Roh · 2023
Closest in time.
Exploring open-vocabulary semantic segmentation from clip vision encoder distillation only
Jun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem, Zhicheng Yan, Chenchen Zhu, Fanyi Xiao, Sean Chang Culatana, and Mohamed Elhoseiny · 2023
Closest in time.
Vision transformers need registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski · 2023
Closest in time.
From images to textual prompts: Zero-shot visual question answering with frozen large language models
Jiaxian Guo, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Boyang Li, Dacheng Tao, and Steven Hoi · 2023
Closest in time.
Global knowledge calibration for fast open-vocabulary segmentation
Kunyang Han, Yong Liu, Jun Hao Liew, Henghui Ding, Jiajun Liu, Yitong Wang, Yansong Tang, Yujiu Yang, Jiashi Feng, Yao Zhao, et al · 2023
Closest in time.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Open-vocabulary semantic segmentation with mask-adapted clip
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu · 2023
Closest in time.
Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation
Huaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He, and Tianrui Li · 2023
Closest in time.
Open vocabulary semantic segmentation with patch aligned contrastive learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang, Ashish Shah, Philip HS Torr, and Ser-Nam Lim · 2023
Closest in time.
Freeseg: Unified, universal and open-vocabulary image segmentation
Jie Qin, Jie Wu, Pengxiang Yan, Ming Li, Ren Yuxi, Xuefeng Xiao, Yitong Wang, Rui Wang, Shilei Wen, Xin Pan, et al · 2023
Closest in time.
Perceptual grouping in contrastive vision-language models
Kanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang, Alexander Toshev, and Jonathon Shlens · 2023
Closest in time.
Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency
Pengzhen Ren, Changlin Li, Hang Xu, Yi Zhu, Guangrun Wang, Jianzhuang Liu, Xiaojun Chang, and Xiaodan Liang · 2023
Closest in time.
Cut and learn for unsupervised object detection and instance segmentation
Xudong Wang, Rohit Girdhar, Stella X Yu, and Ishan Misra · 2023
Closest in time.
A simple baseline for knowledge-based visual question answering
Alexandros Xenos, Themos Stafylakis, Ioannis Patras, and Georgios Tzimiropoulos · 2023
Closest in time.
A simple framework for open-vocabulary segmentation and detection
Hao Zhang, Feng Li, Xueyan Zou, Shilong Liu, Chunyuan Li, Jianwei Yang, and Lei Zhang · 2023
Closest in time.