Fetching the paper…
Reading the bibliography…
Despite the remarkable accuracy of deep neural networks in object detection, they are costly to train and scale due to supervision requirements.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Lsda: Large scale detection through adaptation
Judy Hoffman, Sergio Guadarrama, Eric S Tzeng, Ronghang Hu, Jeff Donahue, Ross Girshick, Trevor Darrell, and Kate Saenko · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Edge boxes: Locating object proposals from edges
C Lawrence Zitnick and Piotr Dollár · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Automatic concept discovery from parallel text and visual corpora
Chen Sun, Chuang Gan, and Ram Nevatia · 2015
Earlier work this paper cites.
Model recommendation: Generating object detectors from few samples
Yu-Xiong Wang and Martial Hebert · 2015
Earlier work this paper cites.
Weakly supervised deep detection networks
Hakan Bilen and Andrea Vedaldi · 2016
Earlier work this paper cites.
Weakly supervised object localization with multi-fold multiple instance learning
Ramazan Gokberk Cinbis, Jakob Verbeek, and Cordelia Schmid · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Large scale semi-supervised object detection using visual and semantic knowledge transfer
Yuxing Tang, Josiah Wang, Boyang Gao, Emmanuel Dellandréa, Robert Gaizauskas, and Liming Chen · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Weakly-supervised visual grounding of phrases with linguistic structures
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee · 2017
Cited alongside, same era.
Zero-shot object detection
Ankan Bansal, Karan Sikka, Gaurav Sharma, Rama Chellappa, and Ajay Divakaran · 2018
Cited alongside, same era.
Knowledge aided consistency for weakly supervised phrase grounding
Kan Chen, Jiyang Gao, and Ram Nevatia · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Transductive learning for zero-shot object detection
Shafin Rahman, Salman Khan, and Nick Barnes · 2019
Later among the works it cites.
C-mil: Continuation multiple instance learning for weakly supervised object detection
Fang Wan, Chang Liu, Wei Ke, Xiangyang Ji, Jianbin Jiao, and Qixiang Ye · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2019
Later among the works it cites.
Cap2det: Learning to amplify weak caption supervision for object detection
Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li, Danfeng Qin, and Jesse Berent · 2019
Later among the works it cites.
Zero shot detection
Pengkai Zhu, Hanxiao Wang, and Venkatesh Saligrama · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
maskrcnn-benchmark: Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch
Francisco Massa and Ross Girshick · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Revisiting knowledge transfer for training object class detectors
Jasper Uijlings, Stefan Popov, and Vittorio Ferrari · 2018
Cited alongside, same era.
Learning to discover and localize visual objects with open vocabulary
Keren Ye, Mingda Zhang, Wei Li, Danfeng Qin, Adriana Kovashka, and Jesse Berent · 2018
Cited alongside, same era.
Multi-level multimodal common semantic space for image-phrase grounding
Hassan Akbari, Svebor Karaman, Surabhi Bhargava, Brian Chen, Carl Vondrick, and Shih-Fu Chang · 2019
Cited alongside, same era.
Learning to detect and retrieve objects from unlabeled videos
Elad Amrani, Rami Ben-Ari, Tal Hakim, and Alex Bronstein · 2019
Cited alongside, same era.
Align2ground: Weakly supervised phrase grounding guided by image-caption alignment
Samyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja, Devi Parikh, and Ajay Divakaran · 2019
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Closest in time.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2020
Closest in time.
A multi-space approach to zero-shot object detection
Dikshant Gupta, Aditya Anantharaman, Nehal Mamgain, Vineeth N Balasubramanian, CV Jawahar, et al · 2020
Closest in time.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Closest in time.
The open images dataset v4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Closest in time.
Weakly-supervised visualbert: Pre-training without parallel images and captions
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, and Kai-Wei Chang · 2020
Closest in time.
Improved visual-semantic alignment for zero-shot object detection
Shafin Rahman, Salman Khan, and Nick Barnes · 2020
Closest in time.
Dlwl: Improving detection for lowshot classes with weakly labelled data
Vignesh Ramanathan, Rui Wang, and Dhruv Mahajan · 2020
Closest in time.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Closest in time.
Don’t even look once: Synthesizing features for zero-shot detection
Pengkai Zhu, Hanxiao Wang, and Venkatesh Saligrama · 2020
Closest in time.