Fetching the paper…
Reading the bibliography…
In computer vision, multi-label recognition are important tasks with many real-world applications, but classifying previously unseen labels remains a significant challenge.
Nus-wide: a real-world web image database from national university of singapore
Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yantao Zheng · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Zero-shot learning by convex combination of semantic embeddings
Mohammad Norouzi, Tomas Mikolov, Samy Bengio, Yoram Singer, Jonathon Shlens, Andrea Frome, Greg S Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Medical image classification with convolutional neural network
Qing Li, Weidong Cai, Xiaogang Wang, Yun Zhou, David Dagan Feng, and Mei Chen · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Label-embedding for image classification
Zeynep Akata, Florent Perronnin, Zaid Harchaoui, and Cordelia Schmid · 2015
Earlier work this paper cites.
Exploit bounding box annotations for multi-label object recognition
Hao Yang, Joey Tianyi Zhou, Yu Zhang, Bin-Bin Gao, Jianxin Wu, and Jianfei Cai · 2016
Earlier work this paper cites.
Yang Zhang, Boqing Gong, and Mubarak Shah · 2016
Earlier work this paper cites.
Multi-label image recognition by recurrently discovering attentional regions
Zhouxia Wang, Tianshui Chen, Guanbin Li, Ruijia Xu, and Liang Lin · 2017
Earlier work this paper cites.
Deep learning in agriculture: A survey
Andreas Kamilaris and Francesc X Prenafeta-Boldú · 2018
Earlier work this paper cites.
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Earlier work this paper cites.
Multi-label zero-shot learning with structured knowledge graphs
Chung-Wei Lee, Wei Fang, Chih-Kuan Yeh, and Yu-Chiang Frank Wang · 2018
Earlier work this paper cites.
Multi-label image classification via knowledge distillation from weakly-supervised detection
Yongcheng Liu, Lu Sheng, Jing Shao, Junjie Yan, Shiming Xiang, and Chunhong Pan · 2018
Earlier work this paper cites.
Image-based manufacturing analytics: Improving the accuracy of an industrial pellet classification system using deep neural networks
Ricardo Rendall, Ivan Castillo, Bo Lu, Brenda Colegrove, Michael Broadway, Leo H Chiang, and Marco S Reis · 2018
Earlier work this paper cites.
Learning semantic-specific graph representation for multi-label image recognition
Tianshui Chen, Muxin Xu, Xiaolu Hui, Hefeng Wu, and Liang Lin · 2019
Earlier work this paper cites.
Multi-label image recognition with graph convolutional networks
Zhao-Min Chen, Xiu-Shen Wei, Peng Wang, and Yanwen Guo · 2019
Cited alongside, same era.
Deep learning-based image recognition for autonomous driving
Hironobu Fujiyoshi, Tsubasa Hirakawa, and Takayoshi Yamashita · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Cited alongside, same era.
Asymmetric loss for multi-label classification
Emanuel Ben-Baruch, Tal Ridnik, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Query2label: A simple transformer way to multi-label classification
Shilong Liu, Lei Zhang, Xiao Yang, Hang Su, and Jun Zhu · 2021
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie · 2021
Later among the works it cites.
Discriminative region-based multi-label zero-shot learning
Sanath Narayan, Akshita Gupta, Salman Khan, Fahad Shahbaz Khan, Ling Shao, and Mubarak Shah · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A shared multi-attention framework for multi-label zero-shot learning
Dat Huynh and Ehsan Elhamifar · 2020
Cited alongside, same era.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang · 2020
Cited alongside, same era.
Multi-label few/zero-shot learning with knowledge aggregated from multiple label graphs
Jueqing Lu, Lan Du, Ming Liu, and Joanna Dipnall · 2020
Cited alongside, same era.
Multi-label graph convolutional network representation learning
Min Shi, Yufei Tang, Xingquan Zhu, and Jianxun Liu · 2020
Cited alongside, same era.
Attention-driven dynamic graph convolutional network for multi-label image recognition
Jin Ye, Junjun He, Xiaojiang Peng, Wenhao Wu, and Yu Qiao · 2020
Cited alongside, same era.
Cross-modality attention with semantic graph embedding for multi-label classification
Renchun You, Zhiyao Guo, Lei Cui, Xiang Long, Yingze Bao, and Shilei Wen · 2020
Cited alongside, same era.
Semantic diversity learning for zero-shot multi-label classification
Avi Ben-Cohen, Nadav Zamir, Emanuel Ben-Baruch, Itamar Friedman, and Lihi Zelnik-Manor · 2021
Cited alongside, same era.
Later among the works it cites.
Imagenet-21k pretraining for the masses
Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor · 2021
Later among the works it cites.
Ml-decoder: Scalable and versatile classification head
Tal Ridnik, Gilad Sharir, Avi Ben-Cohen, Emanuel Ben-Baruch, and Asaf Noy · 2021
Later among the works it cites.
Open-vocabulary object detection using captions
Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang · 2021
Later among the works it cites.
Localized vision-language matching for open-vocabulary object detection
Maria A Bravo, Sudhanshu Mittal, and Thomas Brox · 2022
Closest in time.
Learning to prompt for open-vocabulary object detection with vision-language model
Yu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi, Yue Gao, and Guoqi Li · 2022
Closest in time.
Open vocabulary object detection with pseudo bounding-box labels
Mingfei Gao, Chen Xing, Juan Carlos Niebles, Junnan Li, Ran Xu, Wenhao Liu, and Caiming Xiong · 2022
Closest in time.
Scaling open-vocabulary image segmentation with image-level labels
Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin · 2022
Closest in time.
F-vlm: Open-vocabulary object detection upon frozen vision and language models
Weicheng Kuo, Yin Cui, Xiuye Gu, AJ Piergiovanni, and Anelia Angelova · 2022
Closest in time.
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Closest in time.
Open-vocabulary semantic segmentation with frozen vision-language models
Chaofan Ma, Yuhuan Yang, Yanfeng Wang, Ya Zhang, and Weidi Xie · 2022
Closest in time.
Vision-language pre-training with triple contrastive learning
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang · 2022
Closest in time.