Fetching the paper…
Reading the bibliography…
We present MosaicFusion, a simple yet effective diffusion-based data augmentation approach for large vocabulary instance segmentation.
Otsu N (1979) A threshold selection method from gray-level histograms. TSMC
1979
Earlier work this paper cites.
Miller GA (1995) Wordnet: a lexical database for english. Commun ACM
1995
Earlier work this paper cites.
Di Stefano L, Bulgarelli A (1999) A simple and efficient connected components labeling algorithm. In: ICIAP
1999
Earlier work this paper cites.
Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: CVPR
2009
Earlier work this paper cites.
Kazemzadeh S, Ordonez V, Matten M, Berg T (2014) Referitgame: Referring to objects in photographs of natural scenes. In: EMNLP
2014
Earlier work this paper cites.
Kingma DP, Welling M (2014) Auto-encoding variational bayes. In: ICLR
2014
Earlier work this paper cites.
Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: ECCV
2014
Earlier work this paper cites.
Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. In: MICCAI
2015
Earlier work this paper cites.
Su H, Qi CR, Li Y, Guibas LJ (2015) Render for cnn: Viewpoint estimation in images using cnns trained with rendered 3d model views. In: ICCV
2015
Earlier work this paper cites.
Barron JT, Poole B (2016) The fast bilateral solver. In: ECCV
2016
Earlier work this paper cites.
Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, Franke U, Roth S, Schiele B (2016) The cityscapes dataset for semantic urban scene understanding. In: CVPR
2016
Earlier work this paper cites.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: CVPR
2016
Earlier work this paper cites.
Mao J, Huang J, Toshev A, Camburu O, Yuille AL, Murphy K (2016) Generation and comprehension of unambiguous object descriptions. In: CVPR
2016
Earlier work this paper cites.
Richter SR, Vineet V, Roth S, Koltun V (2016) Playing for data: Ground truth from computer games. In: ECCV
2016
Earlier work this paper cites.
Dwibedi D, Misra I, Hebert M (2017) Cut, paste and learn: Surprisingly easy synthesis for instance detection. In: ICCV
2017
Earlier work this paper cites.
He K, Gkioxari G, Dollár P, Girshick R (2017) Mask r-cnn. In: ICCV
2017
Earlier work this paper cites.
Lin TY, Dollár P, Girshick R, He K, Hariharan B, Belongie S (2017) Feature pyramid networks for object detection. In: CVPR
2017
Earlier work this paper cites.
Loshchilov I, Hutter F (2017) Sgdr: Stochastic gradient descent with warm restarts. In: ICLR
2017
Earlier work this paper cites.
Dvornik N, Mairal J, Schmid C (2018) Modeling visual context is key to augmenting object detection datasets. In: ECCV
2018
Earlier work this paper cites.
Hinterstoisser S, Lepetit V, Wohlhart P, Konolige K (2018) On pre-trained image features and synthetic images for deep learning. In: ECCV Workshops
2018
Earlier work this paper cites.
Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: ACL
2018
Earlier work this paper cites.
Chen K, Wang J, Pang J, Cao Y, Xiong Y, Li X, Sun S, Feng W, Liu Z, Xu J, Zhang Z, Cheng D, Zhu C, Cheng T, Zhao Q, Li B, Lu X, Zhu R, Wu Y, Dai J, Wang J, Shi J, Ouyang W, Loy CC, Lin D (2019) MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:190607155
2019
Earlier work this paper cites.
Fang HS, Sun J, Wang R, Gou M, Li YL, Lu C (2019) Instaboost: Boosting instance segmentation via probability map guided copy-pasting. In: ICCV
2019
Earlier work this paper cites.
Gupta A, Dollar P, Girshick R (2019) Lvis: A dataset for large vocabulary instance segmentation. In: CVPR
2019
Earlier work this paper cites.
Loshchilov I, Hutter F (2019) Decoupled weight decay regularization. In: ICLR
2019
Earlier work this paper cites.
Shao S, Li Z, Zhang T, Peng C, Yu G, Zhang X, Li J, Sun J (2019) Objects365: A large-scale, high-quality dataset for object detection. In: ICCV
2019
Earlier work this paper cites.
Waqas Zamir S, Arora A, Gupta A, Khan S, Sun G, Shahbaz Khan F, Zhu F, Shao L, Xia GS, Bai X (2019) isaid: A large-scale dataset for instance segmentation in aerial images. In: CVPRW
2019
Earlier work this paper cites.
Wu Y, Kirillov A, Massa F, Lo WY, Girshick R (2019) Detectron2
2019
Earlier work this paper cites.
Bochkovskiy A, Wang CY, Liao HYM (2020) Yolov4: Optimal speed and accuracy of object detection. arXiv preprint arXiv:200410934
2020
Earlier work this paper cites.
Han J, Niu M, Du Z, Wei L, Xie L, Zhang X, Tian Q (2020) Joint coco and lvis workshop at eccv 2020: Lvis challenge track technical report: Asynchronous semi-supervised learning for large vocabulary instance segmentation. In: ECCVW
2020
Cited alongside, same era.
Hu X, Jiang Y, Tang K, Chen J, Miao C, Zhang H (2020) Learning to segment the tail. In: CVPR
2020
Cited alongside, same era.
Kuznetsova A, Rom H, Alldrin N, Uijlings J, Krasin I, Pont-Tuset J, Kamali S, Popov S, Malloci M, Kolesnikov A, et al (2020) The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV
2020
Cited alongside, same era.
Li Y, Wang T, Kang B, Tang S, Wang C, Li J, Feng J (2020) Overcoming classifier imbalance for long-tail object detection with balanced group softmax. In: CVPR
2020
Cited alongside, same era.
Ramesh A, Dhariwal P, Nichol A, Chu C, Chen M (2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:220406125
2022
Later among the works it cites.
Rasheed H, Maaz M, Khattak MU, Khan S, Khan FS (2022) Bridging the gap between object and image-level representations for open-vocabulary detection. In: NeurIPS
2022
Later among the works it cites.
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B (2022) High-resolution image synthesis with latent diffusion models. In: CVPR
2022
Later among the works it cites.
Saharia C, Chan W, Saxena S, Li L, Whang J, Denton E, Ghasemipour SKS, Ayan BK, Mahdavi SS, Lopes RG, et al (2022) Photorealistic text-to-image diffusion models with deep language understanding. In: NeurIPS
2022
Later among the works it cites.
Xie J, Zhan X, Liu Z, Ong YS, Loy CC (2022) Delving into inter-image invariance for unsupervised visual representations. IJCV
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu J, Sun Y, Han C, Dou Z, Li W (2020) Deep representation learning on long-tailed data: A learnable embedding augmentation perspective. In: CVPR
2020
Cited alongside, same era.
Ren J, Yu C, Ma X, Zhao H, Yi S, et al (2020) Balanced meta-softmax for long-tailed visual recognition. In: NeurIPS
2020
Cited alongside, same era.
Tan J, Zhang G, Deng H, Wang C, Lu L, Li Q, Dai J (2020b) 1st place solution of lvis challenge 2020: A good box is not a guarantee of a good mask. arXiv preprint arXiv:200901559
2020
Cited alongside, same era.
Wang T, Li Y, Kang B, Li J, Liew J, Tang S, Hoi S, Feng J (2020) The devil is in classification: A simple framework for long-tail instance segmentation. In: ECCV
2020
Cited alongside, same era.
Wu J, Song L, Wang T, Zhang Q, Yuan J (2020) Forest r-cnn: Large-vocabulary long-tailed object detection and instance segmentation. In: ACM-MM
2020
Cited alongside, same era.
Zhan X, Xie J, Liu Z, Ong YS, Loy CC (2020) Online deep clustering for unsupervised representation learning. In: CVPR
2020
Cited alongside, same era.
Ghiasi G, Cui Y, Srinivas A, Qian R, Lin TY, Cubuk ED, Le QV, Zoph B (2021) Simple copy-paste is a strong data augmentation method for instance segmentation. In: CVPR
2021
Cited alongside, same era.
Jia C, Yang Y, Xia Y, Chen YT, Parekh Z, Pham H, Le Q, Sung YH, Li Z, Duerig T (2021) Scaling up visual and vision-language representation learning with noisy text supervision. In: ICML
2021
Cited alongside, same era.
2022
Later among the works it cites.
Zang Y, Li W, Zhou K, Huang C, Loy CC (2022) Open-vocabulary detr with conditional matching. ECCV
2022
Later among the works it cites.
Zhong Y, Yang J, Zhang P, Li C, Codella N, Li LH, Zhou L, Dai X, Yuan L, Li Y, et al (2022) Regionclip: Region-based language-image pretraining. In: CVPR
2022
Later among the works it cites.
Chefer H, Alaluf Y, Vinker Y, Wolf L, Cohen-Or D (2023) Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM Transactions on Graphics (TOG) 42(4):1–10
2023
Closest in time.
Gal R, Arar M, Atzmon Y, Bermano AH, Chechik G, Cohen-Or D (2023) Designing an encoder for fast personalization of text-to-image models. arXiv preprint arXiv:230212228
2023
Closest in time.
Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D (2023) Prompt-to-prompt image editing with cross attention control. In: ICLR
2023
Closest in time.
Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo WY, et al (2023) Segment anything. In: ICCV
2023
Closest in time.
Kuo W, Cui Y, Gu X, Piergiovanni A, Angelova A (2023) F-vlm: Open-vocabulary object detection upon frozen vision and language models. In: ICLR
2023
Closest in time.
Li Z, Zhou Q, Zhang X, Zhang Y, Wang Y, Xie W (2023) Guiding text-to-image diffusion model towards grounded generation. arXiv preprint arXiv:230105221
2023
Closest in time.
OpenAI (2023) Gpt-4 technical report. arXiv preprint arXiv:230308774
2023
Closest in time.
Parmar G, Singh KK, Zhang R, Li Y, Lu J, Zhu JY (2023) Zero-shot image-to-image translation. arXiv preprint arXiv:230203027
2023
Closest in time.
Phung Q, Ge S, Huang JB (2023) Grounded text-to-image synthesis with attention refocusing. arXiv preprint arXiv:230605427
2023
Closest in time.
Podell D, English Z, Lacey K, Blattmann A, Dockhorn T, Müller J, Penna J, Rombach R (2023) Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:230701952
2023
Closest in time.
Wang J, Zhang P, Chu T, Cao Y, Zhou Y, Wu T, Wang B, He C, Lin D (2023) V3det: Vast vocabulary visual detection dataset. In: ICCV
2023
Closest in time.
Wu J, Li X, Xu S, Yuan H, Ding H, Yang Y, Li X, Zhang J, Tong Y, Jiang X, Ghanem B, Tao D (2023) Towards open vocabulary learning: A survey. arXiv preprint arXiv:230615880
2023
Closest in time.
Xie J, Li Y, Huang Y, Liu H, Zhang W, Zheng Y, Shou MZ (2023) Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion. In: ICCV
2023
Closest in time.
Xu S, Li X, Wu S, Zhang W, Li Y, Cheng G, Tong Y, Chen K, Loy CC (2023) Dst-det: Simple dynamic self-training for open-vocabulary object detection. arXiv preprint arXiv:231001393
2023
Closest in time.
Zhang J, Huang J, Jin S, Lu S (2023) Vision-language models for vision tasks: A survey. arXiv preprint arXiv:230400685
2023
Closest in time.
Zhao H, Sheng D, Bao J, Chen D, Chen D, Wen F, Yuan L, Liu C, Zhou W, Chu Q, Zhang W, Yu N (2023) X-paste: Revisiting scalable copy-paste for instance segmentation using clip and stablediffusion. In: ICML
2023
Closest in time.
Zong Z, Song G, Liu Y (2023) Detrs with collaborative hybrid assignments training. In: ICCV
2023
Closest in time.
Lai X, Tian Z, Chen Y, Li Y, Yuan Y, Liu S, Jia J (2024) Lisa: Reasoning segmentation via large language model. In: CVPR
2024
Closest in time.
Lu H, Liu W, Zhang B, Wang B, Dong K, Liu B, Sun J, Ren T, Li Z, Sun Y, et al (2024) Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:240305525
2024
Closest in time.
Rasheed H, Maaz M, Shaji S, Shaker A, Khan S, Cholakkal H, Anwer RM, Xing E, Yang MH, Khan FS (2024) Glamm: Pixel grounding large multimodal model. In: CVPR
2024
Closest in time.
Wu S, Jin S, Zhang W, Xu L, Liu W, Li W, Loy CC (2024) F-lmm: Grounding frozen large multimodal models. arXiv preprint arXiv:240605821
2024
Closest in time.
Zang Y, Li W, Han J, Zhou K, Loy CC (2024) Contextual object detection with multimodal large language models. IJCV
2024
Closest in time.