Fetching the paper…
Reading the bibliography…
The rapid development of Multi-modality Large Language Models (MLLMs) has significantly influenced various aspects of industry and daily life, showcasing impressive capabilities in visual perception and understanding.
N. Murray, L. Marchesotti, and F. Perronnin, “Ava: A large-scale database for aesthetic visual analysis,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2012, pp. 2408–2415
2012
Earlier work this paper cites.
S. Kong, X. Shen, Z. Lin, R. Mech, and C. Fowlkes, “Photo aesthetics ranking network with attributes and content adaptation,” in Proceedings of the European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 662–679
2016
Earlier work this paper cites.
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 586–595
2018
Earlier work this paper cites.
V. Hosu, H. Lin, T. Sziranyi, and D. Saupe, “Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment,” IEEE Transactions on Image Processing (TIP) , vol. 29, pp. 4041–4056, 2020
2020
Earlier work this paper cites.
Y. Fang, H. Zhu, Y. Zeng, K. Ma, and Z. Wang, “Perceptual quality assessment of smartphone photography,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 3677–3686
2020
Earlier work this paper cites.
S. Su, V. Hosu, H. Lin, Y. Zhang, and D. Saupe, “Koniq++: Boosting no-reference image quality assessment in the wild by jointly predicting image quality and defects,” in The 32nd British Machine Vision Conference , 2021
2021
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. N. Team et al. , “Introducing mpt-7b: A new standard for open-source, commercially usable llms,” 2023
2023
Earlier work this paper cites.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” 2023
2023
Earlier work this paper cites.
B. Li, P. Qi, B. Liu, S. Di, J. Liu, J. Pei, J. Yi, and B. Zhou, “Trustworthy ai: From principles to practices,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–46, 2023
2023
Earlier work this paper cites.
F. Liu, K. Lin, L. Li, J. Wang, Y. Yacoob, and L. Wang, “Mitigating hallucination in large multi-modal models via robust instruction tuning,” in The Twelfth International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Towards explainable in-the-wild video quality assessment: a database and a language-prompted approach,” in Proceedings of the 31st ACM International Conference on Multimedia (ACM MM) , 2023, pp. 1045–1054
2023
Cited alongside, same era.
C. Li, Z. Zhang, H. Wu, W. Sun, X. Min, X. Liu, G. Zhai, and W. Lin, “Agiqa-3k: An open database for ai-generated image quality assessment,” IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , 2023
2023
Cited alongside, same era.
H. Wu, E. Zhang, L. Liao, C. Chen, J. Hou, A. Wang, W. Sun, Q. Yan, and W. Lin, “Exploring video quality assessment on user generated contents from aesthetic and technical perspectives,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2023, pp. 20 144–20 154
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong, “Imagereward: Learning and evaluating human preferences for text-to-image generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
H. Liu, C. Li, Y. Li, and Y. J. Lee, “Improved baselines with visual instruction tuning,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
R. Bavishi, E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Taşırlar. (2023) Introducing our multimodal models. [Online]. Available: https://www.adept.ai/blog/fuyu-8b
2023
Cited alongside, same era.
OpenAI, “Gpt-4 technical report,” 2023
2023
Cited alongside, same era.
Google. (2023) Gemini pro. [Online]. Available: https://deepmind.google/technologies/gemini
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang, “mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 13 040–13 051
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
I. Team, “Infimm: Advancing multimodal understanding from flamingo’s legacy through diverse llm integration,” 2024. [Online]. Available: https://huggingface.co/Infi-MM/
2024
Closest in time.
Q. Sun, Y. Cui, X. Zhang, F. Zhang, Q. Yu, Y. Wang, Y. Rao, J. Liu, T. Huang, and X. Wang, “Generative multimodal models are in-context learners,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2024, pp. 14 398–14 409
2024
Closest in time.