Fetching the paper…
Reading the bibliography…
Video saliency prediction aims to identify the regions in a video that attract human attention and gaze, driven by bottom-up features from the video and top-down processes like memory and cognition.
Simple vs complex temporal recurrences for video saliency prediction
Linardos, P.; Mohedano, E.; Nieto, J. J.; O’Connor, N. E.; Giro-i Nieto, X.; and McGuinness, K. 2019 · 1907
Earlier work this paper cites.
How people look at pictures: a study of the psychology and perception in art
Buswell, G. T. 1935 · 1935
Earlier work this paper cites.
Cognitive determinants of fixation location during picture viewing
Loftus, G. R.; and Mackworth, N. H. 1978 · 1978
Earlier work this paper cites.
A model of saliency-based visual attention for rapid scene analysis
Itti, L.; Koch, C.; and Niebur, E. 1998 · 1998
Earlier work this paper cites.
The effects of semantic consistency on eye movements during complex scene viewing
Henderson, J. M.; Weeks Jr, P. A.; and Hollingworth, A. 1999 · 1999
Earlier work this paper cites.
Human gaze control during real-world scene perception
Henderson, J. M. 2003 · 2003
Earlier work this paper cites.
Graph-Based Visual Saliency , 545–552
Schölkopf, B.; Platt, J.; and Hofmann, T. 2007 · 2007
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Eye movements and vision
Yarbus, A. L. 2013 · 2013
Earlier work this paper cites.
SALICON: Reducing the Semantic Gap in Saliency Prediction by Adapting Deep Neural Networks
Huang, X.; Shen, C.; Boix, X.; and Zhao, Q. 2015 · 2015
Earlier work this paper cites.
Actions in the Eye: Dynamic Gaze Datasets and Learnt Saliency Models for Visual Recognition
Mathe, S.; and Sminchisescu, C. 2015 · 2015
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; et al. 2017 · 2017
Earlier work this paper cites.
Dynamic Whitening Saliency
Leborán, V.; García-Díaz, A.; Fdez-Vidal, X. R.; and Pardo, X. M. 2017 · 2017
Earlier work this paper cites.
Salgan: Visual saliency prediction with generative adversarial networks
Pan, J.; Ferrer, C. C.; McGuinness, K.; O’Connor, N. E.; Torres, J.; Sayrol, E.; and Giro-i Nieto, X. 2017 · 2017
Earlier work this paper cites.
Predicting Human Eye Fixations via an LSTM-Based Saliency Attentive Model
Cornia, M.; Baraldi, L.; Serra, G.; and Cucchiara, R. 2018 · 2018
Earlier work this paper cites.
Revisiting salient object detection: Simultaneous detection, ranking, and subitizing of multiple salient objects
Islam, M. A.; Kalash, M.; and Bruce, N. D. 2018 · 2018
Earlier work this paper cites.
Revisiting video saliency: A large-scale benchmark and a new model
Wang, W.; Shen, J.; Guo, F.; Cheng, M.-M.; and Borji, A. 2018 · 2018
Earlier work this paper cites.
Recognize anything: A strong image tagging model
Zhang, Y.; Huang, X.; Ma, J.; Li, Z.; Luo, Z.; Xie, Y.; Qin, Y.; Luo, T.; Li, Y.; Liu, S.; et al. 2024 · 2018
Earlier work this paper cites.
Relative saliency and ranking: Models, metrics, data and benchmarks
Kalash, M.; Islam, M. A.; and Bruce, N. D. 2019 · 2019
Cited alongside, same era.
Tased-net: Temporally-aggregating spatial encoder-decoder network for video saliency detection
Min, K.; and Corso, J. J. 2019 · 2019
Cited alongside, same era.
CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection
Zhang, L.; Zhang, J.; Lin, Z.; Lu, H.; and He, Y. 2019 · 2019
Cited alongside, same era.
Unified image and video saliency modeling
Droste, R.; Jiao, J.; and Noble, J. A. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
DeepVS2.0: A Saliency-Structured Deep Learning Method for Predicting Dynamic Visual Attention
Jiang, L.; Xu, M.; Wang, Z.; and Sigal, L. 2020 · 2020
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Later among the works it cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Li, C.; Yang, J.; Su, H.; Zhu, J.; et al. 2023 · 2023
Later among the works it cites.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z.; Wang, W.; Dong, L.; Hao, Y.; Huang, S.; Ma, S.; and Wei, F. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Video Saliency Prediction Using Spatiotemporal Residual Attentive Networks
Lai, Q.; Wang, W.; Sun, H.; and Shen, J. 2020 · 2020
Cited alongside, same era.
Learning to predict salient faces: A novel visual-audio saliency model
Liu, Y.; Qiao, M.; Xu, M.; Li, B.; Hu, W.; and Borji, A. 2020 · 2020
Cited alongside, same era.
Stavis: Spatio-temporal audiovisual saliency network
Tsiami, A.; Koutras, P.; and Maragos, P. 2020 · 2020
Cited alongside, same era.
Segdiff: Image segmentation with diffusion probabilistic models
Amit, T.; Shaharbany, T.; Nachmani, E.; and Wolf, L. 2021 · 2021
Cited alongside, same era.
Hierarchical domain-adapted feature learning for video saliency prediction
Bellitto, G.; Proietto Salanitri, F.; Palazzo, S.; Rundo, F.; Giordano, D.; and Spampinato, C. 2021 · 2021
Cited alongside, same era.
Video saliency prediction using enhanced spatiotemporal alignment network
Chen, J.; Song, H.; Zhang, K.; Liu, B.; and Liu, Q. 2021 · 2021
Cited alongside, same era.
Video understanding with large language models: A survey
Tang, Y.; Bi, J.; Xu, S.; Song, L.; Liang, S.; Wang, T.; Zhang, D.; An, J.; Lin, J.; Zhu, R.; et al. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
Caption anything: Interactive image description with diverse multimodal controls
Wang, T.; Zhang, J.; Fei, J.; Zheng, H.; Tang, Y.; Li, Z.; Gao, M.; and Zhao, S. 2023 · 2023
Later among the works it cites.
NPF-200: A Multi-Modal Eye Fixation Dataset and Method for Non-Photorealistic Videos
Yang, Z.; Ren, S.; Wu, Z.; Zhao, N.; Wang, J.; Qin, J.; and He, S. 2023 · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Later among the works it cites.
Text-image alignment for diffusion-based perception
Kondapaneni, N.; Marks, M.; Knott, M.; Guimaraes, R.; and Perona, P. 2024 · 2024
Closest in time.
Vila: On pre-training for visual language models
Lin, J.; Yin, H.; Ping, W.; Molchanov, P.; Shoeybi, M.; and Han, S. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Tang, Y.; Shimada, D.; Bi, J.; and Xu, C. 2024 · 2024
Closest in time.
Saliency Prediction on Mobile Videos: A Fixation Mapping-Based Dataset and A Transformer Approach
Wen, S.; Yang, L.; Xu, M.; Qiao, M.; Xu, T.; and Bai, L. 2024 · 2024
Closest in time.
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
Xiong, J.; Zhang, P.; You, T.; Li, C.; Huang, W.; and Zha, Y. 2024 · 2024
Closest in time.
Pink: Unveiling the power of referential comprehension for multi-modal llms
Xuan, S.; Guo, Q.; Yang, M.; and Zhang, S. 2024 · 2024
Closest in time.