Fetching the paper…
Reading the bibliography…
Combining multiple datasets enables performance boost on many computer vision tasks.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
Vision-based traffic sign detection and analysis for intelligent driver assistance systems: Perspectives and survey
A. Mogelmose, M. M. Trivedi, and T. B. Moeslund · 2012
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Fast r-cnn
R. Girshick · 2015
Earlier work this paper cites.
Multiview rgb-d dataset for object instance detection
G. Georgakis, M. A. Reza, A. Mousavian, P.-H. Le, and J. Košecká · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2016
Earlier work this paper cites.
Wider face: A face detection benchmark
S. Yang, P. Luo, C.-C. Loy, and X. Tang · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Cross-domain weakly-supervised object detection through progressive domain adaptation
N. Inoue, R. Furuta, T. Yamasaki, and K. Aizawa · 2018
Cited alongside, same era.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2018
Cited alongside, same era.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Cited alongside, same era.
R-fcn-3000 at 30fps: Decoupling detection and classification
B. Singh, H. Li, A. Sharma, and L. S. Davis · 2018
Cited alongside, same era.
Dota: A large-scale dataset for object detection in aerial images
G.-S. Xia, X. Bai, J. Ding, Z. Zhu, S. Belongie, J. Luo, M. Datcu, M. Pelillo, and L. Zhang · 2018
Borderdet: Border feature for dense object detection
H. Qiu, Y. Ma, Z. Li, S. Liu, and J. Sun · 2020
Later among the works it cites.
Object detection with a unified label space from multiple datasets
X. Zhao, S. Schulter, G. Sharma, Y.-H. Tsai, M. Chandraker, and Y. Wu · 2020
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Later among the works it cites.
Dynamic detr: End-to-end object detection with dynamic attention
X. Dai, Y. Chen, J. Yang, P. Zhang, L. Yuan, and L. Zhang · 2021
Later among the works it cites.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deeplesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning
K. Yan, X. Wang, L. Lu, and R. M. Summers · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Cited alongside, same era.
Fcos: Fully convolutional one-stage object detection
Z. Tian, C. Shen, H. Chen, and T. He · 2019
Cited alongside, same era.
Towards universal object detection by domain attention
X. Wang, Z. Cai, D. Gao, and N. Vasconcelos · 2019
Cited alongside, same era.
Detecting 11k classes: Large scale object detection without fine-grained bounding boxes
H. Yang, H.-Y. Wu, and H. Chen · 2019
Cited alongside, same era.
Later among the works it cites.
Conditional detr for fast training convergence
D. Meng, X. Chen, Z. Fan, G. Zeng, H. Li, Y. Yuan, L. Sun, and J. Wang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Later among the works it cites.
Sparse r-cnn: End-to-end object detection with learnable proposals
P. Sun, R. Zhang, Y. Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang, et al · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
P. Zhang, X. Li, X. Hu, J. Yang, L. Zhang, L. Wang, Y. Choi, and J. Gao · 2021
Later among the works it cites.
Grounded language-image pre-training
L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.-N. Hwang, et al · 2022
Closest in time.
Adavit: Adaptive vision transformers for efficient image recognition
L. Meng, H. Li, B.-C. Chen, S. Lan, Z. Wu, Y.-G. Jiang, and S.-N. Lim · 2022
Closest in time.
Bevt: Bert pretraining of video transformers
R. Wang, D. Chen, Z. Wu, Y. Chen, X. Dai, M. Liu, Y.-G. Jiang, L. Zhou, and L. Yuan · 2022
Closest in time.
Anchor detr: Query design for transformer-based detector
Y. Wang, X. Zhang, T. Yang, and J. Sun · 2022
Closest in time.
Regionclip: Region-based language-image pretraining
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li, et al · 2022
Closest in time.
Simple multi-dataset detection
X. Zhou, V. Koltun, and P. Krähenbühl · 2022
Closest in time.
Vision transformers are good mask auto-labelers
S. Lan, X. Yang, Z. Yu, Z. Wu, J. M. Alvarez, and A. Anandkumar · 2023
Closest in time.