Fetching the paper…
Reading the bibliography…
The Segment Anything Model (SAM), developed by Meta AI Research, represents a significant breakthrough in computer vision, offering a robust framework for image and video segmentation.
J. R. Hobbs, “Granularity,” in Readings in qualitative reasoning about physical systems . Elsevier, 1990, pp. 542–545
1990
Earlier work this paper cites.
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The kitti dataset,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1231–1237, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Song, S. P. Lichtenberg, and J. Xiao, “Sun rgb-d: A rgb-d scene understanding benchmark suite,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 567–576
2015
Earlier work this paper cites.
I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in ICLR , 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE transactions on pattern analysis and machine intelligence , vol. 40, no. 4, pp. 834–848, 2017
2017
Earlier work this paper cites.
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 5828–5839
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 609–617
2017
Earlier work this paper cites.
A. Kurakin, I. Goodfellow, and S. Bengio, “Adversarial machine learning at scale,” in ICLR , 2017
2017
Earlier work this paper cites.
X. Liang, K. Gong, X. Shen, and L. Lin, “Look into person: Joint body parsing & pose estimation network and a new benchmark,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 4, pp. 871–885, 2018
2018
Earlier work this paper cites.
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in ICLR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Bergmann, M. Fauser, D. Sattlegger, and C. Steger, “Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly detection,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 9592–9600
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision . Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
U. Michieli, E. Borsato, L. Rossi, and P. Zanuttigh, “Gmnet: Graph matching network for large scale part semantic segmentation in the wild,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16 . Springer, 2020, pp. 397–414
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 11 621–11 631
2020
Earlier work this paper cites.
M. J. Page, J. E. McKenzie, P. M. Bossuyt, I. Boutron, T. C. Hoffmann, C. D. Mulrow, L. Shamseer, J. M. Tetzlaff, E. A. Akl, S. E. Brennan et al. , “The prisma 2020 statement: an updated guideline for reporting systematic reviews,” bmj , vol. 372, 2021
2021
Earlier work this paper cites.
B. Cheng, A. Schwing, and A. Kirillov, “Per-pixel classification is not all you need for semantic segmentation,” Advances in neural information processing systems , vol. 34, pp. 17 864–17 875, 2021
2021
Earlier work this paper cites.
X. Li, S. Xu, Y. Yang, G. Cheng, Y. Tong, and D. Tao, “Panoptic-partformer: Learning a unified model for panoptic part segmentation,” in European Conference on Computer Vision . Springer, 2022, pp. 729–747
2022
Earlier work this paper cites.
H. K. Cheng and A. G. Schwing, “Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,” 2022
2022
Earlier work this paper cites.
J. Zhou, J. Wang, J. Zhang, W. Sun, J. Zhang, S. Birchfield, D. Guo, L. Kong, M. Wang, and Y. Zhong, “Audio–visual segmentation,” in European Conference on Computer Vision . Springer, 2022, pp. 386–403
2022
Earlier work this paper cites.
Y. Zou, J. Jeong, L. Pemula, D. Zhang, and O. Dabeer, “Spot-the-difference self-supervised pre-training for anomaly detection and segmentation,” in European Conference on Computer Vision . Springer, 2022, pp. 392–408
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
P. Ganugula, Y. Kumar, N. K. Reddy, P. Chellingi, A. K. Thakur, N. Kasera, and C. S. Anand, “Mosaic: Multi-object segmented arbitrary stylization using clip,” 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) , pp. 892–903, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:262464521
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Delatolas, V. S. Kalogeiton, and D. P. Papadopoulos, “Learning the what and how of annotation in video object segmentation,” 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pp. 6936–6946, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:265050493
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
D. Zhang, C. Li, R. Zhang, S. Xie, W. Xue, X. Xie, and S. Zhang, “Fm-ov3d: Foundation model-based cross-modal knowledge blending for open-vocabulary 3d detection,” in AAAI Conference on Artificial Intelligence , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:266520928
2023
Earlier work this paper cites.
IDEA-Research, “Grounded segment anything,” 2023, gitHub repository. [Online]. Available: https://github.com/IDEA-Research/Grounded-Segment-Anything
2023
Earlier work this paper cites.
L. Kong, Y. Liu, X. Li, R. Chen, W. Zhang, J. Ren, L. Pan, K. Chen, and Z. Liu, “Robo3d: Towards robust and reliable 3d perception against corruptions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 19 994–20 006
2023
Earlier work this paper cites.
H. Zhang, F. Li, and N. Ahuja, “Open-nerf: Towards open vocabulary nerf decomposition,” 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pp. 3444–3453, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:264451574
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Liu, Y. Wang, C. Ju, C. Ma, Y. Zhang, and W. Xie, “Annotation-free audio-visual segmentation,” 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) , pp. 5592–5602, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:258762555
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Yang, C. Neary, and U. Topcu, “Multimodal pretrained models for verifiable sequential decision-making: Planning, grounding, and perception,” in Adaptive Agents and Multi-Agent Systems , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:260775783
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
L.-H. Lin, Y. Cui, Y. Hao, F. Xia, and D. Sadigh, “Gesture-informed robot assistance via foundation models,” in Conference on Robot Learning , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:261557194
2023
Cited alongside, same era.
C. Tang, D. Huang, W. Ge, W. Liu, and H. Zhang, “Graspgpt: Leveraging semantic knowledge from a large language model for task-oriented grasping,” IEEE Robotics and Automation Letters , vol. 8, pp. 7551–7558, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:260154903
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Zhou, Z. Wu, R. Boutteau, F. Yang, and D. Ginhac, “Dsec-mos: Segment any moving object with moving ego vehicle,” 2023
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Y. Zhang and X. Zhao, “Mesa: Matching everything by segmenting anything,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 20 217–20 226
2024
Closest in time.
A. E. Saer, L. Grammatikopoulos, G. Sfikas, G. E. Karras, and E. Petsa, “A novel framework for image matching and stitching for moving car inspection under illumination challenges,” Sensors (Basel, Switzerland) , vol. 24, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267851924
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao, “Depth anything: Unleashing the power of large-scale unlabeled data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 10 371–10 381
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
C. Liu, K. Wang, J. Shi, Z. Qiao, and S. Shen, “Fm-fusion: Instance-aware semantic mapping boosted by vision-language foundation models,” IEEE Robotics and Automation Letters , vol. 9, pp. 2232–2239, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267190485
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Sun, H. Zhou, L. Nanbo, L. Chen, J. Zhu, and R. B. Fisher, “A robust deformable linear object perception pipeline in 3d: From segmentation to reconstruction,” IEEE Robotics and Automation Letters , vol. 9, pp. 843–850, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:265653013
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Zhou, F. Liang, L. Chen, H. Liu, Q. Song, G. Vivone, and J. Chanussot, “Mesam: Multiscale enhanced segment anything model for optical remote sensing images,” IEEE Transactions on Geoscience and Remote Sensing , vol. 62, pp. 1–15, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:269670033
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Moghimi, M. Welzel, T. Celik, and T. Schlurmann, “A comparative performance analysis of popular deep learning models and segment anything model (sam) for river water segmentation in close-range remote sensing imagery,” IEEE Access , vol. 12, pp. 52 067–52 085, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:268977332
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Li, T. Chen, X. Wang, Y. Zhong, and X. Xiao, “Adapting the segment anything model for multi-modal retinal anomaly detection and localization,” Information Fusion , p. 102631, 2024
2024
Closest in time.
2024
Closest in time.