Fetching the paper…
Reading the bibliography…
The increased resolution of real-world videos presents a dilemma between efficiency and accuracy for deep Video Quality Assessment (VQA).
R. Keys, “Cubic convolution interpolation for digital image processing,” IEEE Trans. Acoust. Speech Signal Process. , vol. 29, no. 6, pp. 1153–1160, 1981
1981
Earlier work this paper cites.
G. K. Wallace, “The jpeg still picture compression standard,” Commun. ACM , vol. 34, no. 4, p. 30–44, apr 1991
1991
Earlier work this paper cites.
D. J. Le Gall, “The mpeg video compression algorithm,” Signal Processing: Image Communication , vol. 4, no. 2, pp. 129–140, 1992
1992
Earlier work this paper cites.
T. Wiegand, “Draft itu-t recommendation and final draft international standard of joint video specification,” 2003
2003
Earlier work this paper cites.
M. Tagliasacchi, A. Trapanese, S. Tubaro, J. Ascenso, C. Brites, and F. Pereira, “Exploiting spatial redundancy in pixel domain wyner-ziv video coding,” in Proc. IEEE Conf. Image Process. IEEE, 2006, pp. 253–256
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2009, pp. 248–255
2009
Earlier work this paper cites.
A. K. Moorthy and A. C. Bovik, “Blind image quality assessment: From natural scene statistics to perceptual quality,” IEEE Trans. Image Process. , vol. 20, pp. 3350–3364, 2011
2011
Earlier work this paper cites.
M. A. Saad, A. C. Bovik, and C. Charrier, “Blind image quality assessment: A natural scene statistics approach in the dct domain,” IEEE Trans. Image Process. , vol. 21, no. 8, pp. 3339–3352, 2012
2012
Earlier work this paper cites.
A. Mittal, A. K. Moorthy, and A. C. Bovik, “No-reference image quality assessment in the spatial domain,” IEEE Trans. Image Process. , vol. 21, no. 12, pp. 4695–4708, 2012
2012
Earlier work this paper cites.
A. Mittal, R. Soundararajan, and A. C. Bovik, “Making a “completely blind” image quality analyzer,” IEEE Signal Processing Letters , vol. 20, no. 3, pp. 209–212, 2013
2013
Earlier work this paper cites.
R. Soundararajan and A. C. Bovik, “Video quality assessment by reduced reference spatio-temporal entropic differencing,” IEEE Trans. Circuits Syst. Video Technol. , vol. 23, pp. 684–694, 2013
2013
Earlier work this paper cites.
L. Kang, P. Ye, Y. Li, and D. Doermann, “Convolutional neural networks for no-reference image quality assessment.” Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2014
2014
Earlier work this paper cites.
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in EMNLP 2014 . ACL, 2014, pp. 1724–1734
2014
Earlier work this paper cites.
R. A. Poldrack and M. J. Farah, “Progress and challenges in probing the human brain,” Nature , vol. 526, no. 7573, pp. 371–379, 2015
2015
Earlier work this paper cites.
Kang, L. and Ye, P. and Li, Y. and Doermann, D., “Simultaneous estimation of image quality and distortion via multi-task convolutional neural networks.” Proc. IEEE Conf. Image Process. , 2015
2015
Earlier work this paper cites.
F. C. Heilbron, V. Escorcia, B. Ghanem, and J. C. Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2015
2015
Earlier work this paper cites.
A. Mittal, M. A. Saad, and A. C. Bovik, “A completely blind video integrity oracle,” IEEE Trans. Image Process. , vol. 25, no. 1, pp. 289–300, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2016, pp. 770–778
2016
Earlier work this paper cites.
M. Nuutinen, T. Virtanen, M. Vaahteranoksa, T. Vuori, P. Oittinen, and J. Häkkinen, “Cvd2014—a database for evaluating no-reference video quality assessment algorithms,” IEEE Trans. Image Process. , vol. 25, no. 7, pp. 3073–3086, 2016
2016
Earlier work this paper cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in Proc. AAAI Conf. Artif. Intell. , 2017, p. 4278–4284
2017
Earlier work this paper cites.
K. Hara, H. Kataoka, and Y. Satoh, “Learning spatio-temporal features with 3d residual networks for action recognition,” in Proc. Int. Conf. Comput. Vis. Workshops , 2017, pp. 3154–3160
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2017
2017
Earlier work this paper cites.
D. Ghadiyaram and A. C. Bovik, “Perceptual quality prediction on authentically distorted images using a bag of features approach,” Journal of Vision , vol. 17, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. , 2017, p. 6000–6010
2017
Cited alongside, same era.
W. Kay and et al., “The kinetics human action video dataset,” ArXiv , vol. abs/1705.06950, 2017
2017
Cited alongside, same era.
V. Hosu, F. Hahn, M. Jenadeleh, H. Lin, H. Men, T. Szirányi, S. Li, and D. Saupe, “The konstanz natural video database (konvid-1k),” in Ninth International Conference on Quality of Multimedia Experience (QoMEX) , 2017, pp. 1–6
2017
Cited alongside, same era.
R. Goyal, et al., “The ”something something” video database for learning and evaluating visual common sense,” in Proc. Int. Conf. Comput. Vis. , 2017
2017
Cited alongside, same era.
Z. Tu, Y. Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Ugc-vqa: Benchmarking blind video quality assessment for user generated content,” IEEE Trans. Image Process. , 2021
2021
Later among the works it cites.
D. Li and T. Jiang and M. Jiang, “Unified quality assessment of in-the-wild videos with mixed datasets training,” Int. J. Comput. Vis. , vol. 129, no. 4, pp. 1238–1257, 2021
2021
Later among the works it cites.
Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch-vq: ’patching up’ the video quality problem,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 14 019–14 029
2021
Later among the works it cites.
F. Götz-Hahn, V. Hosu, H. Lin, and D. Saupe, “Konvid-150k: A dataset for no-reference video quality assessment of videos in-the-wild,” in IEEE Access 9 . IEEE, 2021, pp. 72 139–72 160
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Buckler, P. Bedoukian, S. Jayasuriya, and A. Sampson, “Eva 2 : Exploiting temporal redundancy in live computer vision,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2018, pp. 533–546
2018
Cited alongside, same era.
W. Kim, J. Kim, S. Ahn, J. Kim, and S. Lee, “Deep video quality assessor: From spatio-temporal visual sensitivity to a convolutional neural aggregation network,” in Proc. Eur. Conf. Comput. Vis. , 2018
2018
Cited alongside, same era.
D. Ghadiyaram, J. Pan, A. C. Bovik, A. K. Moorthy, P. Panda, and K.-C. Yang, “In-capture mobile video distortions: A study of subjective behavior and objective algorithms,” IEEE Trans. Circuits Syst. Video Technol. , vol. 28, no. 9, pp. 2061–2077, 2018
2018
Cited alongside, same era.
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik, “Ava: A video dataset of spatio-temporally localized atomic visual actions,” in CVPR , 2018
2018
Cited alongside, same era.
J. Korhonen, “Two-level approach for no-reference consumer video quality assessment,” IEEE Trans. Image Process. , vol. 28, no. 12, pp. 5923–5938, 2019
2019
Cited alongside, same era.
D. Li, T. Jiang, and M. Jiang, “Quality assessment of in-the-wild videos,” in Proc. ACM Int. Conf. Multimedia , ser. MM ’19, 2019, p. 2351–2359
2019
Cited alongside, same era.
J. You and J. Korhonen, “Deep neural networks for no-reference video quality assessment,” in Proc. IEEE Conf. Image Process. , 2019, pp. 2349–2353
2019
Cited alongside, same era.
D. Li, T. Jiang, W. Lin, and M. Jiang, “Which has better visual quality: The clear blue sky or a blurry animal?” IEEE Trans. Multim. , vol. 21, no. 5, pp. 1221–1234, 2019
2019
Cited alongside, same era.
J. You, “Long short-term convolutional transformer for no-reference video quality assessment,” in Proc. ACM Int. Conf. Multimedia , ser. MM ’21, 2021, p. 2112–2120
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in Proc. Int. Conf. Comput. Vis. , October 2021, pp. 6824–6835
2021
Later among the works it cites.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lucic, and C. Schmid, “Vivit: A video vision transformer,” in Proc. Int. Conf. Comput. Vis. , 2021, pp. 6836–6846
2021
Later among the works it cites.
P. C. Madhusudana, N. Birkbeck, Y. Wang, B. Adsumilli, and A. C. Bovik, “ST-GREED: Space-time generalized entropic differences for frame rate dependent video quality prediction,” IEEE Trans. Image Process. , 2021
2021
Later among the works it cites.
J. P. Ebenezer, Z. Shang, Y. Wu, H. Wei, S. Sethuraman, and A. C. Bovik, “ChipQA: No-reference video quality prediction via space-time chips,” IEEE Trans. Image Process. , vol. 30, pp. 8059–8074, 2021
2021
Later among the works it cites.
B. Chen, L. Zhu, G. Li, F. Lu, H. Fan, and S. Wang, “Learning generalized spatial-temporal deep feature representation for no-reference video quality assessment,” IEEE Trans. Circuits Syst. Video Technol. , 2021
2021
Later among the works it cites.
Y. Liu, X. Zhou, H. Yin, H. Wang, and C. C. Yan, “Efficient video quality assessment with deeper spatiotemporal feature extraction and integration,” Journal of Electronic Imaging , 2021
2021
Later among the works it cites.
A. Kolesnikov and et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
Z. Tu, X. Yu, Y. Wang, N. Birkbeck, B. Adsumilli, and A. C. Bovik, “Rapique: Rapid and accurate video quality prediction of user generated content,” IEEE Open Journal of Signal Processing , vol. 2, pp. 425–440, 2021
2021
Later among the works it cites.
Y. Wang, J. Ke, H. Talebi, J. G. Yim, N. Birkbeck, B. Adsumilli, P. Milanfar, and F. Yang, “Rich features for perceptual quality assessment of ugc videos,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 13 435–13 444
2021
Later among the works it cites.
H. Wang, G. Li, S. Liu, and C.-C. J. Kuo, “Icme 2021 ugc-vqa challenge.” [Online]. Available: http://ugcvqa.com/
2021
Later among the works it cites.
2021
Later among the works it cites.
B. Li, W. Zhang, M. Tian, G. Zhai, and X. Wang, “Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception,” IEEE Trans. Circuits Syst. Video Technol. , vol. 32, no. 9, pp. 5944–5958, 2022
2022
Closest in time.
H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, Q. Yan, and W. Lin, “Discovqa: Temporal distortion-content transformers for video quality assessment,” arXiv preprint arXiv: 2206.09853 , 2022
2022
Closest in time.
H. Wu, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,” Proc. Eur. Conf. Comput. Vis. , 2022
2022
Closest in time.
L. Liao, K. Xu, H. Wu, C. Chen, W. Sun, Q. Yan, and W. Lin, “Exploring the effectiveness of video perceptual representation in blind video quality assessment,” in Proc. ACM Int. Conf. Multimedia , 2022
2022
Closest in time.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 11 976–11 986
2022
Closest in time.