Fetching the paper…

AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation · Around