Fetching the paper…
Reading the bibliography…
The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework.
Making a “completely blind” image quality analyzer
Anish Mittal, Rajiv Soundararajan, and Alan C Bovik · 2012
Earlier work this paper cites.
Blind prediction of natural video quality
Michele A Saad, Alan C Bovik, and Christophe Charrier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
A completely blind video integrity oracle
Anish Mittal, Michele A Saad, and Alan C Bovik · 2015
Earlier work this paper cites.
A quality-of-experience index for streaming video
Zhengfang Duanmu, Kai Zeng, Kede Ma, Abdul Rehman, and Zhou Wang · 2016
Earlier work this paper cites.
Cvd2014—a database for evaluating no-reference video quality assessment algorithms
Mikko Nuutinen, Toni Virtanen, Mikko Vaahteranoksa, Tero Vuori, Pirkko Oittinen, and Jukka Häkkinen · 2016
Earlier work this paper cites.
Study of temporal effects on subjective video quality of experience
Christos George Bampis, Zhi Li, Anush Krishna Moorthy, Ioannis Katsavounidis, Anne Aaron, and Alan Conrad Bovik · 2017
Earlier work this paper cites.
Quality-of-experience of adaptive video streaming: Exploring the space of adaptations
Zhengfang Duanmu, Kede Ma, and Zhou Wang · 2017
Earlier work this paper cites.
The konstanz natural video database (konvid-1k)
Vlad Hosu, Franz Hahn, Mohsen Jenadeleh, Hanhe Lin, Hui Men, Tamás Szirányi, Shujun Li, and Dietmar Saupe · 2017
Earlier work this paper cites.
A survey of rate adaptation techniques for dynamic adaptive streaming over http
Jonathan Kua, Grenville Armitage, and Philip Branch · 2017
Earlier work this paper cites.
Feature-based prediction of streaming video qoe: Distortions, stalling and memory
Christos G Bampis and Alan C Bovik · 2018
Earlier work this paper cites.
A quality-of-experience database for adaptive video streaming
Zhengfang Duanmu, Abdul Rehman, and Zhou Wang · 2018
Earlier work this paper cites.
End-to-end blind quality assessment of compressed videos using deep neural networks
Wentao Liu, Zhengfang Duanmu, and Zhou Wang · 2018
Earlier work this paper cites.
Large-scale study of perceptual video quality
Zeina Sinno and Alan Conrad Bovik · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Quality assessment of in-the-wild videos
Dingquan Li, Tingting Jiang, and Ming Jiang · 2019
Earlier work this paper cites.
Kadid-10k: A large-scale artificially distorted iqa database
Hanhe Lin, Vlad Hosu, and Dietmar Saupe · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Youtube ugc dataset for video compression research
Yilin Wang, Sasi Inguva, and Balu Adsumilli · 2019
Cited alongside, same era.
The waterloo streaming quality-of-experience database-iv
Zhengfang Duanmu, Wentao Liu, Zhuoran Li, Diqi Chen, Zhou Wang, Yizhou Wang, and Wen Gao · 2020
Cited alongside, same era.
Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment
Vlad Hosu, Hanhe Lin, Tamas Sziranyi, and Dietmar Saupe · 2020
Cited alongside, same era.
Towards perceptually optimized adaptive video streaming-a realistic quality of experience database
Christos G Bampis, Zhi Li, Ioannis Katsavounidis, Te-Yuan Huang, Chaitanya Ekanadham, and Alan C Bovik · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Depicting beyond scores: Advancing image quality assessment through multi-modal language models
Zhiyuan You, Zheyuan Li, Jinjin Gu, Zhenfei Yin, Tianfan Xue, and Chao Dong · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al · 2024
Closest in time.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al · 2024
Closest in time.
Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Rich features for perceptual quality assessment of ugc videos
Yilin Wang, Junjie Ke, Hossein Talebi, Joong Gon Yim, Neil Birkbeck, Balu Adsumilli, Peyman Milanfar, and Feng Yang · 2021
Cited alongside, same era.
Patch-vq:’patching up’the video quality problem
Zhenqiang Ying, Maniratnam Mandal, Deepti Ghadiyaram, and Alan Bovik · 2021
Cited alongside, same era.
Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception
Bowen Li, Weixia Zhang, Meng Tian, Guangtao Zhai, and Xianpei Wang · 2022
Cited alongside, same era.
A deep learning based no-reference quality assessment model for ugc videos
Wei Sun, Xiongkuo Min, Wei Lu, and Guangtao Zhai · 2022
Cited alongside, same era.
Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling
Haoning Wu, Chaofeng Chen, Jingwen Hou, Liang Liao, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, et al · 2024
Closest in time.
Aesexpert: Towards multi-modality foundation model for image aesthetics perception
Yipo Huang, Xiangfei Sheng, Zhichao Yang, Quan Yuan, Zhichao Duan, Pengfei Chen, Leida Li, Weisi Lin, and Guangming Shi · 2024
Closest in time.
Continuous and overall quality of experience evaluation for streaming video based on rich features exploration and dual-stage attention
Ziheng Jia, Xiongkuo Min, Wei Sun, and Guangtao Zhai · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Perceptual video quality assessment: A survey
Xiongkuo Min, Huiyu Duan, Wei Sun, Yucheng Zhu, and Guangtao Zhai · 2024
Closest in time.
Analysis of video quality datasets via design of minimalistic video quality models
Wei Sun, Wen Wen, Xiongkuo Min, Long Lan, Guangtao Zhai, and Kede Ma · 2024
Closest in time.
Modular blind video quality assessment
Wen Wen, Mu Li, Yabin Zhang, Yiting Liao, Junlin Li, Li Zhang, and Kede Ma · 2024
Closest in time.
Q-instruct: Improving low-level visual abilities for multi-modality foundation models
Haoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Annan Wang, Kaixin Xu, Chunyi Li, Jingwen Hou, Guangtao Zhai, et al · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Mlvu: A comprehensive benchmark for multi-task long video understanding
Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu · 2024
Closest in time.
Towards open-ended visual quality comparison
Haoning Wu, Hanwei Zhu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Annan Wang, Wenxiu Sun, Qiong Yan, et al · 2025
Closest in time.
In-capture mobile video distortions: A study of subjective behavior and objective algorithms
Deepti Ghadiyaram, Janice Pan, Alan C Bovik, Anush Krishna Moorthy, Prasanjit Panda, and Kai-Chieh Yang · 2077
Closest in time.