Efficient graph-based image segmentation
Felzenszwalb, P. F. and Huttenlocher, D. P · 2004
Earlier work this paper cites.
The pascal visual object classes (VOC) challenge
Everingham, M., Gool, L. V., Williams, C. K. I., Winn, J. M., and Zisserman, A · 2010
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Mottaghi, R., Chen, X., Liu, X., Cho, N., Lee, S., Fidler, S., Urtasun, R., and Yuille, A. L · 2014
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Long, J., Shelhamer, E., and Darrell, T · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Original
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
YFCC100M: the new data in multimedia research
Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Jang, E., Gu, S., and Poole, B · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Pyramid scene parsing network
Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Earlier work this paper cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Zero-shot semantic segmentation
Bucher, M., Vu, T., Cord, M., and Pérez, P · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Miech, A., Zhukov, D., Alayrac, J., Tapaswi, M., Laptev, I., and Sivic, J · 2019
Earlier work this paper cites.