Fetching the paper…
Reading the bibliography…
Copy-Paste is a simple and effective data augmentation strategy for instance segmentation.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Render for cnn: Viewpoint estimation in images using cnns trained with rendered 3d model views
Su, H., Qi, C. R., Li, Y., and Guibas, L. J · 2015
Earlier work this paper cites.
Instance-aware semantic segmentation via multi-task network cascades
Dai, J., He, K., and Sun, J · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Playing for data: Ground truth from computer games
Richter, S. R., Vineet, V., Roth, S., and Koltun, V · 2016
Earlier work this paper cites.
Stc: A simple to complex framework for weakly-supervised semantic segmentation
Wei, Y., Liang, X., Chen, Y., Shen, X., Cheng, M.-M., Feng, J., Zhao, Y., and Yan, S · 2016
Earlier work this paper cites.
Cut, paste and learn: Surprisingly easy synthesis for instance detection
Dwibedi, D., Misra, I., and Hebert, M · 2017
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Weakly supervised semantic segmentation using web-crawled videos
Hong, S., Yeo, D., Kwak, S., Lee, H., and Han, B · 2017
Earlier work this paper cites.
Webly supervised semantic segmentation
Jin, B., Ortiz Segovia, M. V., and Susstrunk, S · 2017
Earlier work this paper cites.
Playing for benchmarks
Richter, S. R., Hayder, Z., and Koltun, V · 2017
Earlier work this paper cites.
Modeling visual context is key to augmenting object detection datasets
Dvornik, N., Mairal, J., and Schmid, C · 2018
Earlier work this paper cites.
On pre-trained image features and synthetic images for deep learning
Hinterstoisser, S., Lepetit, V., Wohlhart, P., and Konolige, K · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and Van Der Maaten, L · 2018
Earlier work this paper cites.
Bootstrapping the performance of webly supervised semantic segmentation
Shen, T., Lin, G., Shen, C., and Reid, I · 2018
Earlier work this paper cites.
Instaboost: Boosting instance segmentation via probability map guided copy-pasting
Fang, H.-S., Sun, J., Wang, R., Gou, M., Li, Y.-L., and Lu, C · 2019
Earlier work this paper cites.
LVIS: A dataset for large vocabulary instance segmentation
Gupta, A., Dollar, P., and Girshick, R · 2019
Earlier work this paper cites.
Photorealistic image synthesis for object instance detection
Hodaň, T., Vineet, V., Gal, R., Shalev, E., Hanzelka, J., Connell, T., Urbina, P., Sinha, S. N., and Guenter, B · 2019
Earlier work this paper cites.
Detectron2
Wu, Y., Kirillov, A., Massa, F., Lo, W.-Y., and Girshick, R · 2019
Earlier work this paper cites.
A survey on instance segmentation: state of the art
Hafiz, A. M. and Bhat, G. M · 2020
Cited alongside, same era.
U2-net: Going deeper with nested u-structure for salient object detection
Qin, X., Zhang, Z., Huang, C., Dehghan, M., Zaiane, O. R., and Jagersand, M · 2020
Cited alongside, same era.
1st place solution of lvis challenge 2020: A good box is not a guarantee of a good mask
Tan, J., Zhang, G., Deng, H., Wang, C., Lu, L., Li, Q., and Dai, J · 2020
Cited alongside, same era.
The devil is in classification: A simple framework for long-tail instance segmentation
Wang, T., Li, Y., Kang, B., Li, J., Liew, J., Tang, S., Hoi, S., and Feng, J · 2020
Cited alongside, same era.
Dynamic head: Unifying object detection heads with attentions
Dai, X., Chen, Y., Xiao, B., Chen, D., Liu, M., Yuan, L., and Zhang, L · 2021
Cited alongside, same era.
Promptdet: Expand your detector vocabulary with uncurated images
Feng, C., Zhong, Y., Jie, Z., Chu, X., Ren, H., Wei, X., Xie, W., and Ma, L · 2022
Closest in time.
Vector quantized diffusion model for text-to-image synthesis
Gu, S., Chen, D., Bao, J., Wen, F., Zhang, B., Chen, D., Yuan, L., and Guo, B · 2022
Closest in time.
Mask dino: Towards a unified transformer-based framework for object detection and segmentation
Li, F., Zhang, H., Liu, S., Zhang, L., Ni, L. M., Shum, H.-Y., et al · 2022
Closest in time.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al · 2022
Closest in time.
Image segmentation using text and image prompts
Lüddecke, T. and Ecker, A · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
Gu, X., Lin, T.-Y., Kuo, W., and Cui, Y · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M · 2021
Cited alongside, same era.
On model calibration for long-tailed object detection and instance segmentation
Pan, T.-Y., Zhang, C., Li, Y., Hu, H., Xuan, D., Changpinyo, S., Gong, B., and Chao, W.-L · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Meng, L., Dai, X., Chen, Y., Zhang, P., Chen, D., Liu, M., Wang, J., Wu, Z., Yuan, L., and Jiang, Y.-G · 2022
Closest in time.
Matteformer: Transformer-based image matting via prior-tokens
Park, G., Son, S., Yoo, J., Kim, S., and Kwak, N · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Closest in time.
Bridging the gap between object and image-level representations for open-vocabulary detection
Rasheed, H., Maaz, M., Khattak, M. U., Khan, S., and Khan, F. S · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Closest in time.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al · 2022
Closest in time.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Closest in time.
Su, Y., Deng, J., Sun, R., Lin, G., and Wu, Q · 2022
Closest in time.
Omnivl: One foundation model for image-language and video-language tasks
Wang, J., Chen, D., Wu, Z., Luo, C., Zhou, L., Zhao, Y., Xie, Y., Liu, C., Jiang, Y.-G., and Yuan, L · 2022
Closest in time.
Scaling autoregressive models for content-rich text-to-image generation
Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al · 2022
Closest in time.
Selfreformer: Self-refined network with transformer for salient object detection
Yun, Y. K. and Lin, W · 2022
Closest in time.
Regionclip: Region-based language-image pretraining
Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al · 2022
Closest in time.
Detecting twenty-thousand classes using image-level supervision
Zhou, X., Girdhar, R., Joulin, A., Krähenbühl, P., and Misra, I · 2022
Closest in time.
Frido: Feature pyramid diffusion for complex scene image synthesis
Fan, W.-C., Chen, Y.-C., Chen, D., Cheng, Y., Yuan, L., and Wang, Y.-C. F · 2023
Closest in time.