Fetching the paper…
Reading the bibliography…
The rapid development of Multi-modality Large Language Models (MLLMs) has navigated a paradigm shift in computer vision, moving towards versatile foundational models.
“Recommendation 500-10: Methodology for the subjective assessment of the quality of television pictures,” ITU-R Rec. BT.500, 2000
2000
Earlier work this paper cites.
A. Hore and D. Ziou, “Image quality metrics: Psnr vs. ssim,” in International Conference on Pattern Recognition . IEEE, 2010, pp. 2366–2369
2010
Earlier work this paper cites.
N. Murray, L. Marchesotti, and F. Perronnin, “Ava: A large-scale database for aesthetic visual analysis,” in CVPR , 2012, pp. 2408–2415
2012
Earlier work this paper cites.
D. Jayaraman, A. Mittal, A. K. Moorthy, and A. C. Bovik, “Objective quality assessment of multiply distorted images,” in ASILOMAR , 2012, pp. 1693–1697
2012
Earlier work this paper cites.
A. Mittal, R. Soundararajan, and A. C. Bovik, “Making a “completely blind” image quality analyzer,” IEEE Signal Processing Letters , vol. 20, no. 3, pp. 209–212, 2013
2013
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,” Transactions of the Association for Computational Linguistics , vol. 2, pp. 67–78, 2014
2014
Earlier work this paper cites.
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and C. L. Zitnick, “Microsoft coco captions: Data collection and evaluation server,” 2015
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in ICCV , 2015
2015
Earlier work this paper cites.
D. Ghadiyaram and A. C. Bovik, “Massive online crowdsourced study of subjective and objective picture quality,” IEEE TIP , vol. 25, no. 1, pp. 372–387, 2015
2015
Earlier work this paper cites.
S. Kong, X. Shen, Z. Lin, R. Mech, and C. Fowlkes, “Photo aesthetics ranking network with attributes and content adaptation,” in ECCV , 2016
2016
Earlier work this paper cites.
D. Ghadiyaram and A. C. Bovik, “Massive online crowdsourced study of subjective and objective picture quality,” IEEE , vol. 25, no. 1, pp. 372–387, 2016
2016
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in CVPR , 2018, pp. 586–595
2018
Earlier work this paper cites.
E. Prashnani, H. Cai, Y. Mostofi, and P. Sen, “Pieapp: Perceptual image-error assessment through pairwise preference,” in CVPR , June 2018
2018
Earlier work this paper cites.
H. Talebi and P. Milanfar, “Nima: Neural image assessment,” IEEE TIP , 2018
2018
Earlier work this paper cites.
D. Li, T. Jiang, W. Lin, and M. Jiang, “Which has better visual quality: The clear blue sky or a blurry animal?” IEEE TMM , vol. 21, no. 5, pp. 1221–1234, 2019
2019
Earlier work this paper cites.
H. Lin, V. Hosu, and D. Saupe, “Kadid-10k: A large-scale artificially distorted iqa database,” in QoMEX , 2019, pp. 1–3
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, “Ok-vqa: A visual question answering benchmark requiring external knowledge,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Earlier work this paper cites.
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson, “nocaps: novel object captioning at scale,” in ICCV , 2019
2019
Earlier work this paper cites.
V. Hosu, H. Lin, T. Sziranyi, and D. Saupe, “Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,” IEEE TIP , vol. 29, pp. 4041–4056, 2020
2020
Earlier work this paper cites.
Y. Fang, H. Zhu, Y. Zeng, K. Ma, and Z. Wang, “Perceptual quality assessment of smartphone photography,” in CVPR , 2020
2020
Earlier work this paper cites.
J. Gu, H. Cai, H. Chen, X. Ye, J. Ren, and C. Dong, “Pipal: a large-scale image quality assessment dataset for perceptual image restoration,” 2020
2020
Earlier work this paper cites.
T. Guha, V. Hosu, D. Saupe, B. Goldlücke, N. Kumar, W. Lin, V. Martinez, K. Somandepalli, S. Narayanan, W.-H. Cheng, K. McLaughlin, H. Adam, J. See, and L.-K. Wong, “Atqam/mast’20: Joint workshop on aesthetic and technical quality assessment of multimedia and media analytics for societal trends,” in ACM MM , 2020, p. 4758–4760
2020
Earlier work this paper cites.
Z. Ying, H. Niu, P. Gupta, D. Mahajan, D. Ghadiyaram, and A. Bovik, “From patches to pictures (paq-2-piq): Mapping the perceptual space of picture quality,” in CVPR , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Su, V. Hosu, H. Lin, Y. Zhang, and D. Saupe, “Koniq++ : Boosting no-reference image quality assessment in the wild by jointly predicting image quality and defects,” in The British Machine Vision Conference (BMVC) , 2021, pp. 1–12
2021
Cited alongside, same era.
Y. Wang, J. Ke, H. Talebi, J. G. Yim, N. Birkbeck, B. Adsumilli, P. Milanfar, and F. Yang, “Rich features for perceptual quality assessment of ugc videos,” in CVPR , June 2021, pp. 13 435–13 444
2021
Cited alongside, same era.
Z. Ying, M. Mandal, D. Ghadiyaram, and A. Bovik, “Patch-vq: ’patching up’ the video quality problem,” in CVPR , 2021
2021
Cited alongside, same era.
Z. Zhang, W. Sun, X. Min, W. Zhu, T. Wang, W. Lu, and G. Zhai, “A no-reference evaluation metric for low-light image enhancement,” in IEEE International Conference on Multimedia and Expo . IEEE, 2021, pp. 1–6
2021
Cited alongside, same era.
Google, “Gemini pro,” 2023. [Online]. Available: https://deepmind.google/technologies/gemini
2023
Later among the works it cites.
OpenAI, “Gpt-4 technical report,” 2023
2023
Later among the works it cites.
Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin, “Mmbench: Is your multi-modal model an all-around player?” 2023
2023
Later among the works it cites.
J. Lu, J. Rao, K. Chen, X. Guo, Y. Zhang, B. Sun, C. Yang, and J. Yang, “Evaluation and mitigation of agnosia in multimodal large language models,” 2023
2023
Later among the works it cites.
B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan, “Seed-bench: Benchmarking multimodal llms with generative comprehension,” 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” 2021
2021
Cited alongside, same era.
H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,” in ECCV , 2022
2022
Cited alongside, same era.
Z. Zhang, W. Sun, X. Min, W. Zhu, T. Wang, and G. Zhai, “A no-reference deep learning quality assessment method for super-resolution images based on frequency maps,” in IEEE International Symposium on Circuits and Systems , 2022, pp. 3170–3174
2022
Cited alongside, same era.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 320–335
2022
Cited alongside, same era.
H. Wu, C. Chen, J. Hou, L. Liao, A. Wang, W. Sun, Q. Yan, and W. Lin, “Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling,” in ECCV , 2022
2022
Cited alongside, same era.
J. Wang, K. C. K. Chan, and C. C. Loy, “Exploring clip for assessing the look and feel of images,” 2022
2022
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023
2023
Cited alongside, same era.
M. N. Team. (2023) Introducing mpt-7b: A new standard for open-source, commercially usable llms. Accessed: 2023-05-05. [Online]. Available: www.mosaicml.com/blog/mpt-7b
2023
Cited alongside, same era.
J. Hou, W. Lin, Y. Fang, H. Wu, C. Chen, L. Liao, and W. Liu, “Towards transparent deep image aesthetics assessment with tag-based content descriptors,” IEEE TIP , 2023
2023
Later among the works it cites.
H. Wu, C. Chen, L. Liao, J. Hou, W. Sun, Q. Yan, J. Gu, and W. Lin, “Neighbourhood representative sampling for efficient end-to-end video quality assessment,” 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, C. Jiang, C. Li, Y. Xu, H. Chen, J. Tian, Q. Qi, J. Zhang, and F. Huang, “mplug-owl: Modularization empowers large language models with multimodality,” 2023
2023
Later among the works it cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
2023
Later among the works it cites.
R. Bavishi, E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Taşırlar, “Introducing our multimodal models,” 2023. [Online]. Available: https://www.adept.ai/blog/fuyu-8b
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” 2023
2023
Later among the works it cites.
P. Zhang, X. Dong, B. Wang, Y. Cao, C. Xu, L. Ouyang, Z. Zhao, S. Ding, S. Zhang, H. Duan, W. Zhang, H. Yan, X. Zhang, W. Li, J. Li, K. Chen, C. He, X. Zhang, Y. Qiao, D. Lin, and J. Wang, “Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition,” 2023
2023
Later among the works it cites.
Huggingface, “Introducing idefics: An open reproduction of state-of-the-art visual language model,” 2023. [Online]. Available: https://huggingface.co/blog/idefics
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
SkunkworksAI, “Bakllava,” 2024. [Online]. Available: https://github.com/SkunkworksAI/BakLLaVA
2024
Closest in time.
H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, and W. Lin, “Q-bench: A benchmark for general-purpose foundation models on low-level vision,” in ICLR , 2024
2024
Closest in time.
I. Team, “Infimm: Advancing multimodal understanding from flamingo’s legacy through diverse llm integration,” 2024. [Online]. Available: https://huggingface.co/Infi-MM/
2024
Closest in time.
C. Zhang, S. Su, Y. Zhu, Q. Yan, J. Sun, and Y. Zhang, “Exploring and evaluating image restoration potential in dynamic scenes,” in CVPR , 2022, pp. 2057–2066
2066
Closest in time.