Fetching the paper…
Reading the bibliography…
We introduce the first zero-shot approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models.
Semantic object classes in video: A high-definition ground truth database
Gabriel J. Brostow, Julien Fauqueur, and Roberto Cipolla · 2009
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Semantic video cnns through representation warping, 2017
Raghudeep Gadde, Varun Jampani, and Peter V. Gehler · 2017
Earlier work this paper cites.
Pyramid scene parsing network
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia · 2017
Earlier work this paper cites.
Deep feature flow for video recognition, 2017
Xizhou Zhu, Yuwen Xiong, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Efficient semantic video segmentation with per-frame inference, 2020
Yifan Liu, Chunhua Shen, Changqian Yu, and Jingdong Wang · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Earlier work this paper cites.
Mask2former for video instance segmentation, 2021
Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexander Kirillov, Rohit Girdhar, and Alexander G. Schwing · 2021
Earlier work this paper cites.
Vspw: A large-scale dataset for video scene parsing in the wild
Jiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang, Guangrui Li, and Yi Yang · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
Temporal memory attention for video semantic segmentation
Hao Wang, Weining Wang, and Jing Liu · 2021
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Later among the works it cites.
Perceptual grouping in contrastive vision-language models
K. Ranasinghe, B. McKinzie, S. Ravi, Y. Yang, A. Toshev, and J. Shlens · 2023
Later among the works it cites.
Emergent correspondence from image diffusion, 2023
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan · 2023
Later among the works it cites.
Dvis: Decoupled video instance segmentation framework, 2023
Tao Zhang, Xingye Tian, Yu Wu, Shunping Ji, Xuebo Wang, Yuan Zhang, and Pengfei Wan · 2023
Later among the works it cites.
Dvis++: Improved decoupled framework for universal video segmentation, 2023
Tao Zhang, Xingye Tian, Yikang Zhou, Shunping Ji, Xuebo Wang, Xin Tao, Yuan Zhang, Pengfei Wan, Zhongyuan Wang, and Yu Wu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Cited alongside, same era.
Temporally efficient vision transformer for video instance segmentation
Shusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang, Jiemin Fang, Liu, Xun Zhao, and Ying Shan · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach · 2023
Cited alongside, same era.
Unsupervised keypoints from pretrained diffusion models
Eric Hedlin, Gopal Sharma, Shweta Mahajan, Xingzhe He, Hossam Isack, Abhishek Kar Helge Rhodin, Andrea Tagliasacchi, and Kwang Moo Yi · 2023
Cited alongside, same era.
Diffusion hyperfeatures: Searching through time and space for semantic correspondence, 2023
Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holynski, and Trevor Darrell · 2023
Cited alongside, same era.
Univs: Unified and universal video segmentation with prompts as queries, 2024
Minghan Li, Shuai Li, Xindong Zhang, and Lei Zhang · 2024
Closest in time.
Open-vocabulary attention maps with token optimization for semantic segmentation in diffusion models, 2024
Pablo Marcos-Manchón, Roberto Alcover-Couso, Juan C. SanMiguel, and Jose M. Martínez · 2024
Closest in time.
Emerdiff: Emerging pixel-level semantic knowledge in diffusion models
Koichi Namekata, Amirmojtaba Sabour, Sanja Fidler, and Seung Wook Kim · 2024
Closest in time.
A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang · 2024
Closest in time.