Fetching the paper…
Reading the bibliography…
Object detection is a fundamental task in computer vision, requiring large annotated datasets that are difficult to collect, as annotators need to label objects and their bounding boxes.
Vl-bert: Pre-training of generic visual-linguistic representations
Su, W.; Zhu, X.; Cao, Y.; Li, B.; Lu, L.; Wei, F.; and Dai, J. 2019 · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H.; and Bansal, M. 2019 · 1908
Earlier work this paper cites.
Scene graph parsing by attention graph
Andrews, M.; Chia, Y. K.; and Witteveen, S. 2019 · 1909
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Chen, Y.-C.; Li, L.; Yu, L.; Kholy, A. E.; Ahmed, F.; Gan, Z.; Cheng, Y.; and Liu, J. 2019 · 1909
Earlier work this paper cites.
Solving the multiple instance problem with axis-parallel rectangles
Dietterich, T. G.; Lathrop, R. H.; and Lozano-Pérez, T. 1997 · 1997
Earlier work this paper cites.
Contrastive Learning for Weakly Supervised Phrase Grounding
Gupta, T.; Vahdat, A.; Chechik, G.; Yang, X.; Kautz, J.; and Hoiem, D. 2020 · 2006
Earlier work this paper cites.
Human action recognition in videos using kinematic features and multiple instance learning
Ali, S.; and Shah, M. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010 · 2010
Earlier work this paper cites.
Segmentation as selective search for object recognition
Van de Sande, K. E.; Uijlings, J. R.; Gevers, T.; and Smeulders, A. W. 2011 · 2011
Earlier work this paper cites.
Multiple instance classification: Review, taxonomy and comparative study
Amores, J. 2013 · 2013
Earlier work this paper cites.
Selective search for object recognition
Uijlings, J. R.; Van De Sande, K. E.; Gevers, T.; and Smeulders, A. W. 2013 · 2013
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P.; Lai, A.; Hodosh, M.; and Hockenmaier, J. 2014 · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X.; Fang, H.; Lin, T.-Y.; Vedantam, R.; Gupta, S.; Dollár, P.; and Zitnick, C. L. 2015 · 2015
Earlier work this paper cites.
Is object localization for free?-weakly-supervised learning with convolutional neural networks
Oquab, M.; Bottou, L.; Laptev, I.; and Sivic, J. 2015 · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Schuster, S.; Krishna, R.; Chang, A.; Fei-Fei, L.; and Manning, C. D. 2015 · 2015
Earlier work this paper cites.
Weakly supervised deep detection networks
Bilen, H.; and Vedaldi, A. 2016 · 2016
Earlier work this paper cites.
Audio event detection using weakly labeled data
Kumar, A.; and Raj, B. 2016 · 2016
Cited alongside, same era.
On Support Relations and Semantic Scene Graphs arXiv preprint arXiv:1609.05834
Liao, W.; Yang, M. Y.; Ackermann, H.; and Rosenhahn, B. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Zhou, B.; Khosla, A.; Lapedriza, A.; Oliva, A.; and Torralba, A. 2016 · 2016
Cited alongside, same era.
Mask r-cnn
He, K.; Gkioxari, G.; Dollár, P.; and Girshick, R. 2017 · 2017
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
Hu, R.; Rohrbach, M.; Andreas, J.; Darrell, T.; and Saenko, K. 2017 · 2017
Sniper: Efficient multi-scale training
Singh, B.; Najibi, M.; and Davis, L. S. 2018 · 2018
Later among the works it cites.
Scene graph parsing as dependency parsing
Wang, Y.-S.; Liu, C.; Zeng, X.; and Yuille, A. 2018 · 2018
Later among the works it cites.
Multi-level multimodal common semantic space for image-phrase grounding
Akbari, H.; Karaman, S.; Bhargava, S.; Chen, B.; Vondrick, C.; and Chang, S.-F. 2019 · 2019
Later among the works it cites.
Precise Detection in Densely Packed Scenes
Goldman, E.; Herzig, R.; Eisenschtat, A.; Goldberger, J.; and Hassner, T. 2019 · 2019
Later among the works it cites.
Spatio-temporal action graph networks
Herzig, R.; Levi, E.; Xu, H.; Gao, H.; Brosh, E.; Wang, X.; Globerson, A.; and Darrell, T. 2019 · 2019
Later among the works it cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Cited alongside, same era.
Multiple-instance learning for medical image and video analysis
Quellec, G.; Cazuguel, G.; Cochener, B.; and Lamard, M. 2017 · 2017
Cited alongside, same era.
Multiple instance detection network with online instance classifier refinement
Tang, P.; Wang, X.; Bai, X.; and Liu, W. 2017 · 2017
Cited alongside, same era.
Scene Graph Generation by Iterative Message Passing
Xu, D.; Zhu, Y.; Choy, C. B.; and Fei-Fei, L. 2017 · 2017
Cited alongside, same era.
Neural Motifs: Scene Graph Parsing with Global Context
Zellers, R.; Yatskar, M.; Thomson, S.; and Choi, Y. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Later among the works it cites.
Triplet-Aware Scene Graph Embeddings
Schroeder, B.; Tripathi, S.; and Tang, H. 2019 · 2019
Later among the works it cites.
Scene graph captioner: Image captioning based on structural visual representation
Xu, N.; Liu, A.-A.; Liu, J.; Nie, W.; and Su, Y. 2019 · 2019
Later among the works it cites.
Cap2det: Learning to amplify weak caption supervision for object detection
Ye, K.; Zhang, M.; Kovashka, A.; Li, W.; Qin, D.; and Berent, J. 2019 · 2019
Later among the works it cites.
Grounded video description
Zhou, L.; Kalantidis, Y.; Chen, X.; Corso, J. J.; and Rohrbach, M. 2019 · 2019
Later among the works it cites.
Compositional Video Synthesis with Action Graphs
Bar, A.; Herzig, R.; Wang, X.; Chechik, G.; Darrell, T.; and Globerson, A. 2020 · 2020
Closest in time.
Learning Canonical Representations for Scene Graph to Image Generation
Herzig, R.; Bar, A.; Xu, H.; Chechik, G.; Darrell, T.; and Globerson, A. 2020 · 2020
Closest in time.
Something-Else: Compositional Action Recognition with Spatial-Temporal Interaction Networks
Materzynska, J.; Xiao, T.; Herzig, R.; Xu, H.; Wang, X.; and Darrell, T. 2020 · 2020
Closest in time.
Differentiable Scene Graphs
Raboh, M.; Herzig, R.; Chechik, G.; Berant, J.; and Globerson, A. 2020 · 2020
Closest in time.
Instance-Aware, Context-Focused, and Memory-Efficient Weakly Supervised Object Detection
Ren, Z.; Yu, Z.; Yang, X.; Liu, M.-Y.; Lee, Y. J.; Schwing, A. G.; and Kautz, J. 2020 · 2020
Closest in time.