Fetching the paper…
Reading the bibliography…
The realm of computer vision has witnessed a paradigm shift with the advent of foundational models, mirroring the transformative influence of large language models in the domain of natural language processing.
Brown T, Mann B, Ryder N, et al (2020) Language models are few-shot learners. Adv Neural Inform Process Syst 33:1877–1901
1901
Earlier work this paper cites.
Gu Z, Zhou S, Niu L, et al (2020) Context-aware feature generation for zero-shot semantic segmentation. In: Proceedings of the 28th ACM International Conference on Multimedia, pp 1921–1929
1929
Earlier work this paper cites.
2007
Earlier work this paper cites.
Everingham M, Gool LV, Williams CKI, et al (2010) The pascal visual object classes (VOC) challenge. Int J Comput Vis 88(2):303–338
2010
Earlier work this paper cites.
Mottaghi R, Chen X, Liu X, et al (2014) The role of context for object detection and semantic segmentation in the wild. In: Conf. Comput. Vis. Pattern Recog. IEEE Computer Society, pp 891–898
2014
Earlier work this paper cites.
Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: Conf. Comput. Vis. Pattern Recog., pp 3431–3440
2015
Earlier work this paper cites.
Chen LC, Papandreou G, Kokkinos I, et al (2017) Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans Pattern Anal Mach Intell 40(4):834–848
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Caesar H, Uijlings JRR, Ferrari V (2018) Coco-stuff: Thing and stuff classes in context. In: Conf. Comput. Vis. Pattern Recog. Computer Vision Foundation / IEEE Computer Society, pp 1209–1218
2018
Earlier work this paper cites.
Devlin J, Chang MW, Lee K, et al (2018) Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:181004805
2018
Earlier work this paper cites.
Radford A, Narasimhan K, Salimans T, et al (2018) Improving language understanding by generative pre-training. OpenAI
2018
Earlier work this paper cites.
Angus M, Czarnecki K, Salay R (2019) Efficacy of pixel-level ood detection for semantic segmentation. arXiv preprint arXiv:191102897
2019
Earlier work this paper cites.
Bucher M, Vu TH, Cord M, et al (2019) Zero-shot semantic segmentation. Advances in Neural Information Processing Systems 32
2019
Earlier work this paper cites.
Gupta A, Dollár P, Girshick RB (2019) LVIS: A dataset for large vocabulary instance segmentation. In: CVPR. Computer Vision Foundation / IEEE, pp 5356–5364
2019
Earlier work this paper cites.
Lu J, Batra D, Parikh D, et al (2019) Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Adv Neural Inform Process Syst 32
2019
Earlier work this paper cites.
Nguyen K, Todorovic S (2019) Feature weighting and boosting for few-shot segmentation. In: Int. Conf. Comput. Vis. IEEE, pp 622–631
2019
Earlier work this paper cites.
Radford A, Wu J, Child R, et al (2019) Language models are unsupervised multitask learners. OpenAI blog 1(8):9
2019
Earlier work this paper cites.
Xian Y, Choudhury S, He Y, et al (2019) Semantic projection network for zero-and few-label semantic segmentation. In: Conf. Comput. Vis. Pattern Recog., pp 8256–8265
2019
Earlier work this paper cites.
Cui Z, Longshi W, Wang R (2020) Open set semantic segmentation with statistical test and adaptive threshold. In: 2020 IEEE International Conference on Multimedia and Expo (ICME), IEEE, pp 1–6
2020
Earlier work this paper cites.
Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. Adv Neural Inform Process Syst 33:6840–6851
2020
Cited alongside, same era.
Li X, Wei T, Chen YP, et al (2020) FSS-1000: A 1000-class dataset for few-shot segmentation. In: Conf. Comput. Vis. Pattern Recog. Computer Vision Foundation / IEEE, pp 2866–2875
2020
Cited alongside, same era.
Xia Y, Zhang Y, Liu F, et al (2020) Synthesize then compare: Detecting failures and anomalies for semantic segmentation. In: Eur. Conf. Comput. Vis., Springer, pp 145–161
2020
Cited alongside, same era.
Cen J, Yun P, Cai J, et al (2021) Deep metric learning for open world semantic segmentation. In: Int. Conf. Comput. Vis., pp 15,333–15,342
2021
Cited alongside, same era.
Cheng J, Nandi S, Natarajan P, et al (2021) SIGN: spatial-information incorporated generative network for generalized zero-shot semantic segmentation. In: Int. Conf. Comput. Vis. IEEE, pp 9536–9546
Cheng Y, Li L, Xu Y, et al (2023) Segment and track anything. arXiv preprint arXiv:230506558
2023
Closest in time.
Chowdhery A, Narang S, Devlin J, et al (2023) Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24(240):1–113
2023
Closest in time.
2023
Closest in time.
Hammam A, Bonarens F, Ghobadi SE, et al (2023) Identifying out-of-domain objects with dirichlet deep neural networks. In: Int. Conf. Comput. Vis., pp 4560–4569
2023
Closest in time.
Jiang PT, Yang Y (2023) Segment anything is a good pseudo-label generator for weakly supervised semantic segmentation. arXiv preprint arXiv:230501275
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Radford A, Kim JW, Hallacy C, et al (2021) Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn., Proceedings of Machine Learning Research, vol 139. PMLR, pp 8748–8763
2021
Cited alongside, same era.
Ranftl R, Bochkovskiy A, Koltun V (2021) Vision transformers for dense prediction. In: Conf. Comput. Vis. Pattern Recog., pp 12,179–12,188
2021
Cited alongside, same era.
Ghiasi G, Gu X, Cui Y, et al (2022) Scaling open-vocabulary image segmentation with image-level labels. In: Eur. Conf. Comput. Vis., vol 13696. Springer, pp 540–557
2022
Cited alongside, same era.
He K, Chen X, Xie S, et al (2022) Masked autoencoders are scalable vision learners. In: Conf. Comput. Vis. Pattern Recog., pp 16,000–16,009
2022
Cited alongside, same era.
Li J, Li D, Xiong C, et al (2022) BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In: Int. Conf. Mach. Learn., Proceedings of Machine Learning Research, vol 162. PMLR, pp 12,888–12,900
2022
Cited alongside, same era.
Liu Q, Wen Y, Han J, et al (2022) Open-world semantic segmentation via contrasting and clustering vision-language embedding. In: Eur. Conf. Comput. Vis., Springer, pp 275–292
2022
Cited alongside, same era.
Ma C, Yang Y, Wang Y, et al (2022) Open-vocabulary semantic segmentation with frozen vision-language models. arXiv preprint arXiv:221015138
2022
Cited alongside, same era.
Closest in time.
Kirillov A, Mintun E, Ravi N, et al (2023) Segment anything. arXiv preprint arXiv:230402643
2023
Closest in time.
Liang F, Wu B, Dai X, et al (2023) Open-vocabulary semantic segmentation with mask-adapted CLIP. In: Conf. Comput. Vis. Pattern Recog. IEEE, pp 7061–7070
2023
Closest in time.
Luo H, Bao J, Wu Y, et al (2023) Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation. In: Int. Conf. Mach. Learn., PMLR, pp 23,033–23,044
2023
Closest in time.
Oin J, Wu J, Yan P, et al (2023) Freeseg: Unified, universal and open-vocabulary image segmentation. In: Conf. Comput. Vis. Pattern Recog. IEEE, pp 19,446–19,455
2023
Closest in time.
Oquab M, Darcet T, Théo Moutakanni ea (2023) Dinov2: Learning robust visual features without supervision. CoRR
2023
Closest in time.
Ramanathan V, Kalia A, Petrovic V, et al (2023) PACO: parts and attributes of common objects. In: CVPR. IEEE, pp 7141–7151
2023
Closest in time.
Shen Q, Yang X, Wang X (2023) Anything-3d: Towards single-view anything reconstruction in the wild. arXiv preprint arXiv:230410261
2023
Closest in time.
Tang L, Xiao H, Li B (2023) Can sam segment anything? when sam meets camouflaged object detection. arXiv preprint arXiv:230404709
2023
Closest in time.
Touvron H, Lavril T, Izacard G, et al (2023) Llama: Open and efficient foundation language models. arXiv preprint arXiv:230213971
2023
Closest in time.
Xu M, Zhang Z, Wei F, et al (2023) Side adapter network for open-vocabulary semantic segmentation. In: Conf. Comput. Vis. Pattern Recog., pp 2945–2954
2023
Closest in time.
Yang J, Gao M, Li Z, et al (2023) Track anything: Segment anything meets videos. arXiv preprint arXiv:230411968
2023
Closest in time.
Zhang K, Liu D (2023) Customized segment anything model for medical image segmentation. arXiv preprint arXiv:230413785
2023
Closest in time.
Zhao W, Rao Y, Liu Z, et al (2023) Unleashing text-to-image diffusion models for visual perception. CoRR
2023
Closest in time.
Guo J, Hao Z, Wang C, et al (2024) Data-efficient large vision models through sequential autoregression. arXiv preprint arXiv:240204841
2024
Closest in time.