Fetching the paper…
Reading the bibliography…
As an alternative to expensive expert evaluation, Image Aesthetic Assessment (IAA) stands out as a crucial task in computer vision.
Content-based photo quality assessment
Xiaoou Tang, Wei Luo, and Xiaogang Wang. 2013 · 1943
Earlier work this paper cites.
A model of aesthetic appreciation and aesthetic judgments
Helmut Leder, Benno Belke, Andries Oeberst, and Dorothee Augustin. 2004 · 2004
Earlier work this paper cites.
Prospects for a cognitive neuroscience of visual aesthetics
Anjan Chatterjee. 2004 · 2004
Earlier work this paper cites.
AVA: A large-scale database for aesthetic visual analysis. In 2012 IEEE conference on computer vision and pattern recognition . IEEE, 2408–2415
Naila Murray, Luca Marchesotti, and Florent Perronnin. 2012 · 2012
Earlier work this paper cites.
Methodology for the subjective assessment of the quality of television pictures
B Series. 2012 · 2012
Earlier work this paper cites.
Recognizing Image Style. In Proceedings of the British Machine Vision Conference . BMVA Press
Sergey Karayev, Matthew Trentacoste, Helen Han, Aseem Agarwala, Trevor Darrell, Aaron Hertzmann, and Holger Winnemoeller. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO Captions: Data Collection and Evaluation Server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
VQA: Visual Question Answering. In ICCV
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
ClickSmart: A context-aware viewpoint recommendation system for mobile photography
Yogesh Singh Rawat and Mohan S Kankanhalli. 2016 · 2016
Earlier work this paper cites.
Photo aesthetics ranking network with attributes and content adaptation. In European conference on computer vision . Springer, 662–679
Shu Kong, Xiaohui Shen, Zhe Lin, Radomir Mech, and Charless Fowlkes. 2016 · 2016
Earlier work this paper cites.
Joint image and text representation for aesthetics analysis. In Proceedings of the 24th ACM international conference on Multimedia . 262–266
Ye Zhou, Xin Lu, Junping Zhang, and James Z Wang. 2016 · 2016
Earlier work this paper cites.
Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C. Berg, and Tamara L. Berg. 2016 · 2016
Earlier work this paper cites.
Image aesthetic assessment: An experimental survey
Yubin Deng, Chen Change Loy, and Xiaoou Tang. 2017 · 2017
Earlier work this paper cites.
Multigranular event recognition of personal photo albums
Cong Guo, Xinmei Tian, and Tao Mei. 2017 · 2017
Earlier work this paper cites.
Personalized Image Aesthetics. In The IEEE International Conference on Computer Vision (ICCV)
Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J. Foran. 2017 · 2017
Earlier work this paper cites.
Aesthetic critiques generation for photos. In Proceedings of the IEEE international conference on computer vision . 3514–3523
Kuang-Yu Chang, Kung-Hung Lu, and Chu-Song Chen. 2017 · 2017
Earlier work this paper cites.
Personalized image aesthetics. In Proceedings of the IEEE international conference on computer vision . 638–647
Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J Foran. 2017 · 2017
Earlier work this paper cites.
A-lamp: Adaptive layout-aware multi-patch deep convolutional neural network for photo aesthetic assessment. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4535–4544
Shuang Ma, Jing Liu, and Chang Wen Chen. 2017 · 2017
Earlier work this paper cites.
NIMA: Neural image assessment
Hossein Talebi and Peyman Milanfar. 2018 · 2018
Earlier work this paper cites.
Photographic composition classification and dominant geometric element detection for outdoor scenes
Jun-Tae Lee, Han-Ul Kim, Chul Lee, and Chang-Su Kim. 2018 · 2018
Earlier work this paper cites.
Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In ACL
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Earlier work this paper cites.
A comprehensive survey on image aesthetic quality assessment. In 2019 IEEE/ACIS 18th International Conference on Computer and Information Science (ICIS) . IEEE, 294–299
Hongtao Yang, Ping Shi, Saike He, Da Pan, Zefeng Ying, and Ling Lei. 2019 · 2019
Earlier work this paper cites.
Aesthetic attributes assessment of images. In Proceedings of the 27th ACM international conference on multimedia . 311–319
Xin Jin, Le Wu, Geng Zhao, Xiaodong Li, Xiaokun Zhang, Shiming Ge, Dongqing Zou, Bin Zhou, and Xinghui Zhou. 2019 · 2019
Earlier work this paper cites.
Effective Aesthetics Prediction With Multi-Level Spatially Pooled Features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Vlad Hosu, Bastian Goldlucke, and Dietmar Saupe. 2019 · 2019
Earlier work this paper cites.
Neural aesthetic image reviewer
Wenshan Wang, Su Yang, Weishan Zhang, and Jiulong Zhang. 2019 · 2019
Earlier work this paper cites.
nocaps: novel object captioning at scale. In ICCV
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. 2019 · 2019
Earlier work this paper cites.
OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. In Conference on Computer Vision and Pattern Recognition (CVPR)
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Earlier work this paper cites.
Modeling image composition for visual aesthetic assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops . 0–0
Dong Liu, Rohit Puri, Nagendra Kamath, and Subhabrata Bhattacharya. 2019 · 2019
Earlier work this paper cites.
Roundness-preserving warping for aesthetic enhancement-based stereoscopic image editing
Xiongli Chai, Feng Shao, Qiuping Jiang, and Yo-Sung Ho. 2020 · 2020
Cited alongside, same era.
Adaptive Fractional Dilated Convolution Network for Image Aesthetics Assessment. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Qiuyu Chen, Wei Zhang, Ning Zhou, Peng Lei, Yi Xu, Yu Zheng, and Jianping Fan. 2020 · 2020
Cited alongside, same era.
EVA: An Explainable Visual Aesthetics Dataset. In Joint Workshop on Aesthetic and Technical Quality Assessment of Multimedia and Media Analytics for Societal Trends (ATQAM/MAST’20), ACM Multimedia . ACM, Seattle, United States, 5–13
Chen Kang, Giuseppe Valenzise, and Frédéric Dufaux. 2020 · 2020
Cited alongside, same era.
Infrared and visible cross-modal image retrieval through shared features
Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao, Feng Yang, Anyong Qin, and Deyu Meng. 2021 · 2021
Cited alongside, same era.
VideoChat: Chat-Centric Video Understanding
Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wen Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023 · 2023
Later among the works it cites.
Are emergent abilities of Large Language Models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. 2023 · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a · 2023
Later among the works it cites.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2023b · 2023
Later among the works it cites.
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Cited alongside, same era.
Hierarchical layout-aware graph convolutional network for unified aesthetics assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8475–8484
Dongyu She, Yu-Kun Lai, Gaoxiong Yi, and Kun Xu. 2021 · 2021
Cited alongside, same era.
Distilling knowledge from object classification to aesthetics assessment
Jingwen Hou, Henghui Ding, Weisi Lin, Weide Liu, and Yuming Fang. 2022 · 2022
Cited alongside, same era.
Comment-guided semantics-aware image aesthetics assessment
Yuzhen Niu, Shanshan Chen, Bingrui Song, Zhixian Chen, and Wenxi Liu. 2022 · 2022
Cited alongside, same era.
Personalized Image Aesthetics Assessment With Rich Attributes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 19861–19869
Yuzhe Yang, Liwu Xu, Leida Li, Nan Qie, Yaqian Li, Peng Zhang, and Yandong Guo. 2022 · 2022
Cited alongside, same era.
Rethinking Image Aesthetics Assessment: Models, Datasets and Benchmarks. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , Lud De Raedt (Ed.). International Joint Conferences on Artificial Intelligence Organization, 942–948
Shuai He, Yongchang Zhang, Rui Xie, Dongxiang Jiang, and Anlong Ming. 2022 · 2022
Cited alongside, same era.
Haoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Annan Wang, Chunyi Li, Wenxiu Sun, Qiong Yan, Guangtao Zhai, and Weisi Lin. 2023 · 2023
Later among the works it cites.
Theme-aware Visual Attribute Reasoning for Image Aesthetics Assessment
Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang, Yaqian Li, Yandong Guo, and Guangming Shi. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
Otter: A Multi-Modal Model with In-Context Instruction Tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. 2023 · 2023
Later among the works it cites.
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, Hongsheng Li, and Yu Qiao. 2023 · 2023
Later among the works it cites.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023 · 2023
Later among the works it cites.
Pan Zhang, Xiaoyi Dong, Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Wenwei Zhang, Hang Yan, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. 2023 · 2023
Later among the works it cites.
MMBench: Is Your Multi-modal Model an All-around Player?
Yuanzhan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2023 · 2023
Later among the works it cites.
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Zhenyu Qiu, Wei Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, and Rongrong Ji. 2023 · 2023
Later among the works it cites.
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. 2023 · 2023
Later among the works it cites.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. 2023 · 2023
Later among the works it cites.
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, et al · 2023
Later among the works it cites.
Peng Jin, Ryuichi Takanobu, Caiwan Zhang, Xiaochun Cao, and Li Yuan. 2023 · 2023
Later among the works it cites.
Q-align: Teaching lmms for visual scoring via discrete text-defined levels
Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Yixuan Gao, Annan Wang, Erli Zhang, Wenxiu Sun, et al · 2023
Later among the works it cites.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2023 · 2023
Later among the works it cites.
Q-instruct: Improving low-level visual abilities for multi-modality foundation models
Haoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Annan Wang, Kaixin Xu, Chunyi Li, Jingwen Hou, Guangtao Zhai, et al · 2023
Later among the works it cites.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kenneth KY Wong, Zhenguo Li, and Hengshuang Zhao. 2023 · 2023
Later among the works it cites.
Sharegpt4v: Improving large multi-modal models with better captions
Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. 2023 · 2023
Later among the works it cites.
OpenAI. 2023b · 2023
Later among the works it cites.
Probing Sentiment-Oriented Pre-Training Inspired by Human Sentiment Perception Mechanism. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2850–2860
Tinglei Feng, Jiaxuan Liu, and Jufeng Yang. 2023 · 2023
Later among the works it cites.
EgoPlan-Bench: Benchmarking Egocentric Embodied Planning with Multimodal Large Language Models
Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li, Rui Wang, Ruifeng Xu, Ying Shan, and Xihui Liu. 2023 · 2023
Later among the works it cites.
Towards explainable in-the-wild video quality assessment: a database and a language-prompted approach. In Proceedings of the 31st ACM International Conference on Multimedia . 1045–1054
Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. 2023 · 2023
Later among the works it cites.
Moe-llava: Mixture of experts for large vision-language models
Bin Lin, Zhenyu Tang, Yang Ye, Jiaxi Cui, Bin Zhu, Peng Jin, Junwu Zhang, Munan Ning, and Li Yuan. 2024 · 2024
Closest in time.