Fetching the paper…
Reading the bibliography…
Recent advances in zero-shot image recognition suggest that vision-language models learn generic visual representations with a high degree of semantic information that may be arbitrarily probed with natural language phrases.
Laws of organization in perceptual forms
Max Wertheimer · 1938
Earlier work this paper cites.
Some methods for classification and analysis of multivariate observations
J. MacQueen · 1967
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
David Marr · 1982
Earlier work this paper cites.
Robust analysis of feature spaces: Color image segmentation
Dorin Comaniciu and Peter Meer · 1997
Earlier work this paper cites.
Normalized cuts and image segmentation
Jianbo Shi and Jitendra Malik · 2000
Earlier work this paper cites.
Visual grouping and object recognition
Jitendra Malik · 2001
Earlier work this paper cites.
Learning a classification model for segmentation
Xiaofeng Ren and Jitendra Malik · 2003
Earlier work this paper cites.
Cortical algorithms for perceptual grouping
Pieter R Roelfsema et al · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Philipp Krähenbühl and Vladlen Koltun · 2011
Earlier work this paper cites.
Semantic segmentation using regions and parts
Pablo Arbeláez, Bharath Hariharan, Chunhui Gu, Saurabh Gupta, Lubomir Bourdev, and Jitendra Malik · 2012
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Selective search for object recognition
Jasper RR Uijlings, Koen EA Van De Sande, Theo Gevers, and Arnold WM Smeulders · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
The three R’s of computer vision: Recognition, reconstruction and reorganization
Jitendra Malik, Pablo Arbeláez, João Carreira, Katerina Fragkiadaki, Ross Girshick, Georgia Gkioxari, Saurabh Gupta, Bharath Hariharan, Abhishek Kar, and Shubham Tulsiani · 2016
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Open vocabulary scene parsing
Hang Zhao, Xavier Puig, Bolei Zhou, Sanja Fidler, and Antonio Torralba · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Excessive invariance causes adversarial vulnerability
Jörn-Henrik Jacobsen, Jens Behrmann, Richard Zemel, and Matthias Bethge · 2018
Earlier work this paper cites.
End-to-end joint semantic segmentation of actors and actions in video
Jingwei Ji, Shyamal Buch, Alvaro Soto, and Juan Carlos Niebles · 2018
Earlier work this paper cites.
The book of why: the new science of cause and effect
Judea Pearl and Dana Mackenzie · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
Semantic understanding of scenes through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2018
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
Zero-shot semantic segmentation
Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Pérez · 2019
Earlier work this paper cites.
Jointly discovering visual objects and spoken words from raw sensory input
David F. Harwath, Adrià Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James R. Glass · 2019
Earlier work this paper cites.
Bipartite conditional random fields for panoptic segmentation
Sadeep Jayasumana, Kanchana Ranasinghe, Mayuka Jayawardhana, Sahan Damith Liyanaarachchi, and Harsha Ranasinghe · 2019
Earlier work this paper cites.
Invariant information clustering for unsupervised image classification and segmentation
Xu Ji, Andrea Vedaldi, and João F. Henriques · 2019
Earlier work this paper cites.
Do better imagenet models transfer better?
Simon Kornblith, Jonathon Shlens, and Quoc V Le · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Cited alongside, same era.
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 2019
Cited alongside, same era.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models
Jianmo Ni, Gustavo Hernández Ábrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang · 2021
Later among the works it cites.
A closer look at self-training for zero-label semantic segmentation
Giuseppe Pastore, Fabio Cermelli, Yongqin Xian, Massimiliano Mancini, Zeynep Akata, and Barbara Caputo · 2021
Later among the works it cites.
Open-world entity segmentation
Lu Qi, Jason Kuen, Yi Wang, Jiuxiang Gu, Hengshuang Zhao, Zhe Lin, Philip Torr, and Jiaya Jia · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Semantic projection network for zero-and few-label semantic segmentation
Yongqin Xian, Subhabrata Choudhury, Yang He, Bernt Schiele, and Zeynep Akata · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, T. J. Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeff Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Context-aware feature generation for zero-shot semantic segmentation
Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang · 2020
Cited alongside, same era.
From pixel to patch: Synthesize context-aware features for zero-shot semantic segmentation
Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang · 2020
Cited alongside, same era.
Later among the works it cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Later among the works it cites.
Self-supervised disentangled representation learning for third-person imitation learning
Jinghuan Shang and Michael S. Ryoo · 2021
Later among the works it cites.
Conterfactual generative zero-shot semantic segmentation
Feihong Shen, Jun Liu, and Ping Hu · 2021
Later among the works it cites.
Multi-grained vision language pre-training: Aligning texts with visual concepts
Yan Zeng, Xinsong Zhang, and Hang Li · 2021
Later among the works it cites.
Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai · 2021
Later among the works it cites.
Discovering objects that can move
Zhipeng Bao, Pavel Tokmakov, A. Jabri, Yu-Xiong Wang, Adrien Gaidon, and Martial Hebert · 2022
Closest in time.
Peekaboo: Text to image diffusion models are zero-shot segmentors
Ryan Burgert, Kanchana Ranasinghe, Xiang Li, and Michael S. Ryoo · 2022
Closest in time.
Coyo-700m: Image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim · 2022
Closest in time.
Junbum Cha, Jonghwan Mun, and Byung-Seok Roh · 2022
Closest in time.
Yufeng Cui, Lichen Zhao, Feng Liang, Yangguang Li, and Jing Shao · 2022
Closest in time.
Coarse-to-fine vision-language pre-training with fusion in the backbone
Zi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang, Jianfeng Wang, Linjie Li, Zicheng Liu, Ce Liu, Yann LeCun, Nanyun Peng, Jianfeng Gao, and Lijuan Wang · 2022
Closest in time.
Savi++: Towards end-to-end object-centric learning from real-world videos
Gamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Michael Curtis Mozer, and Thomas Kipf · 2022
Closest in time.
Open-vocabulary image segmentation
Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin · 2022
Closest in time.
Unsupervised semantic segmentation by distilling feature correspondences
Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T Freeman · 2022
Closest in time.
Object discovery and representation networks
Olivier J. H’enaff, Skanda Koppula, Evan Shelhamer, Daniel Zoran, Andrew Jaegle, Andrew Zisserman, João Carreira, and Relja Arandjelovi’c · 2022
Closest in time.
Language-driven semantic segmentation
Boyi Li, Kilian Q Weinberger, Serge Belongie, Vladlen Koltun, and René Ranftl · 2022
Closest in time.
Adapting clip for phrase localization without further training
Jiahao Li, Greg Shakhnarovich, and Raymond A. Yeh · 2022
Closest in time.
Does self-supervised learning really improve reinforcement learning from pixels?
Xiang Li, Jinghuan Shang, Srijan Das, and Michael S. Ryoo · 2022
Closest in time.
Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation
Huaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He, and Tianrui Li · 2022
Closest in time.
Deep spectral methods: A surprisingly strong baseline for unsupervised semantic segmentation and localization
Luke Melas-Kyriazi, C. Rupprecht, Iro Laina, and Andrea Vedaldi · 2022
Closest in time.
A comprehensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes
Mazda Moayeri, Phillip Pope, Yogesh Balaji, and Soheil Feizi · 2022
Closest in time.
Open vocabulary semantic segmentation with patch aligned contrastive learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang, Ashish Shah, Philip H. S. Torr, and Ser Nam Lim · 2022
Closest in time.
Learning viewpoint-agnostic visual representations by recovering tokens in 3d space
Jinghuan Shang, Srijan Das, and Michael S. Ryoo · 2022
Closest in time.
Reclip: A strong zero-shot baseline for referring expression comprehension
Sanjay Subramanian, Will Merrill, Trevor Darrell, Matt Gardner, Sameer Singh, and Anna Rohrbach · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross · 2022
Closest in time.
Self-supervised visual representation learning with semantic grouping
Xin Wen, Bingchen Zhao, Anlin Zheng, X. Zhang, and Xiaojuan Qi · 2022
Closest in time.
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang · 2022
Closest in time.
FILIP: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Patch-level representation learning for self-supervised vision transformers
Sukmin Yun, Hankook Lee, Jaehyung Kim, and Jinwoo Shin · 2022
Closest in time.
Unsupervised semantic segmentation with self-supervised object-centric representations
Andrii Zadaianchuk, Matthaeus Kleindessner, Yi Zhu, Francesco Locatello, and Thomas Brox · 2022
Closest in time.
Position prediction as an effective pretraining strategy
Shuangfei Zhai, Navdeep Jaitly, Jason Ramapuram, Dan Busbridge, Tatiana Likhomanenko, Joseph Y Cheng, Walter Talbott, Chen Huang, Hanlin Goh, and Joshua M Susskind · 2022
Closest in time.
Glipv2: Unifying localization and vision-language understanding
Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen, Liunian Harold Li, Xiyang Dai, Lijuan Wang, Lu Yuan, Jenq-Neng Hwang, and Jianfeng Gao · 2022
Closest in time.
Diffusion models for zero-shot open-vocabulary segmentation
Laurynas Karazija, Iro Laina, Andrea Vedaldi, and C. Rupprecht · 2023
Closest in time.
Scaling language-image pre-training via masking
Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He · 2023
Closest in time.
Learning open-vocabulary semantic segmentation models from natural language supervision
Jilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng, Yi Wang, Yu Qiao, and Weidi Xie · 2023
Closest in time.
Associating spatially-consistent grouping with text-supervised semantic segmentation
Yabo Zhang, Zihao Wang, Jun Hao Liew, Jingjia Huang, Manyu Zhu, Jiashi Feng, and Wangmeng Zuo · 2023
Closest in time.