Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) have become increasingly important due to their state-of-the-art performance and ability to integrate multiple data modalities, such as text, images, and audio, to perform complex tasks with high accuracy.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Automatic attribute discovery and characterization from noisy web data
Tamara L. Berg, Alexander C. Berg, and Jonathan Shih. 2010 · 2010
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Image-based recommendations on styles and substitutes
Julian McAuley, Christopher Targett, Qinfeng Shi, and Anton Van Den Hengel. 2015 · 2015
Earlier work this paper cites.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Ruining He and Julian McAuley. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan Yuille, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
Automatic spatially-aware fashion concept discovery
Xintong Han, Zuxuan Wu, Phoenix X Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S Davis. 2017 · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Preference dynamics with multimodal user-item interactions in social media recommendation
Dimitrios Rafailidis, Pavlos Kefalas, and Yannis Manolopoulos. 2017 · 2017
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
Pog: personalized outfit generation for fashion recommendation at alibaba ifashion
Wen Chen, Pipei Huang, Jiaming Xu, Xin Guo, Cheng Guo, Fei Sun, Chao Li, Andreas Pfadler, Huan Zhao, and Binqiang Zhao. 2019 · 2019
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Earlier work this paper cites.
Fashion iq: A new dataset towards retrieving images by natural language feedback
Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. 2021 · 2021
Earlier work this paper cites.
Personalised clip or: how to find your vacation videos
2022 · 2022
Earlier work this paper cites.
“this is my unicorn, fluffy”: Personalizing frozen vision-language representations
Niv Cohen, Rinon Gal, Eli A Meirom, Gal Chechik, and Yuval Atzmon. 2022 · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Earlier work this paper cites.
Understanding political polarization via jointly modeling users, connections and multimodal contents on heterogeneous graphs
Hanjia Lyu and Jiebo Luo. 2022 · 2022
Earlier work this paper cites.
Multimodal persona based generation of comic dialogs
Harsh Agrawal, Aditya Mishra, Manish Gupta, et al. 2023 · 2023
Earlier work this paper cites.
On task-personalized multimodal few-shot learning for visually-rich document entity retrieval
Jiayi Chen, Hanjun Dai, Bo Dai, Aidong Zhang, and Wei Wei. 2023 · 2023
Earlier work this paper cites.
Athena 3.0: personalized multimodal chatbot with neuro-symbolic dialogue generators
Yue Fan, Kevin K Bowden, Wen Cui, Winson Chen, Vrindavan Harrison, Angela Ramirez, Saaket Agashe, Xinyue Gabby Liu, Neha Pullabhotla, NQJ Bheemanpally, et al. 2023 · 2023
Earlier work this paper cites.
Vitr: augmenting vision transformers with relation-focused learning for cross-modal information retrieval
Yan Gong, Georgina Cosma, and Axel Finke. 2023 · 2023
Cited alongside, same era.
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. 2023 · 2023
Cited alongside, same era.
A content-driven micro-video recommendation dataset at scale
Yongxin Ni, Yu Cheng, Xiangyan Liu, Junchen Fu, Youhua Li, Xiangnan He, Yongfeng Zhang, and Fajie Yuan. 2023 · 2023
Cited alongside, same era.
Scalable multimodal learning and multimedia recommendation
Jialie Shen, Marie Morrison, and Zhu Li. 2023 · 2023
Cited alongside, same era.
User-aware prefix-tuning is a good learner for personalized image captioning
Xuan Wang, Guanhong Wang, Wenhao Chai, Jiayu Zhou, and Gaoang Wang. 2023 · 2023
Cited alongside, same era.
Interarec: Interactive recommendations using multimodal large language models
Saketh Reddy Karra and Theja Tulabandhula. 2024 · 2024
Closest in time.
Layout-and-retouch: A dual-stage framework for improving diversity in personalized image generation
Kangyeol Kim, Wooseok Seo, Sehyun Nam, Bodam Kim, Suhyeon Jeong, Wonwoo Cho, Jaegul Choo, and Youngjae Yu. 2024 · 2024
Closest in time.
Longlamp: A benchmark for personalized long-form text generation
Ishita Kumar, Snigdha Viswanathan, Sushrita Yerra, Alireza Salemi, Ryan A Rossi, Franck Dernoncourt, Hanieh Deilamsalehy, Xiang Chen, Ruiyi Zhang, Shubham Agarwal, et al. 2024 · 2024
Closest in time.
A multimodal generative ai copilot for human pathology
Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Melissa Zhao, Aaron K Chow, Kenji Ikemura, Ahrong Kim, Dimitra Pouli, Ankush Patel, et al. 2024 · 2024
Closest in time.
Llm-rec: Personalized recommendation via prompting large language models
Hanjia Lyu, Song Jiang, Hanqing Zeng, Yinglong Xia, Qifan Wang, Si Zhang, Ren Chen, Chris Leung, Jiajie Tang, and Jiebo Luo. 2024a · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-modal self-supervised learning for recommendation
Wei Wei, Chao Huang, Lianghao Xia, and Chuxu Zhang. 2023 · 2023
Cited alongside, same era.
The dawn of lmms: Preliminary explorations with gpt-4v (ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. 2023 · 2023
Cited alongside, same era.
Cgsmp: Controllable generative summarization via multimodal prompt
Qian Yong, Jueqi Wei, YiRen Zhang, XiLun Zhang, Chao Wei, Simiao Chen, Yunhe Li, Cheng Ye, Bing Huang, and Hao Wang. 2023 · 2023
Cited alongside, same era.
Exploring recommendation capabilities of gpt-4v (ision): A preliminary case study
Peilin Zhou, Meng Cao, You-Liang Huang, Qichen Ye, Peiyan Zhang, Junling Liu, Yueqi Xie, Yining Hua, and Jaeboum Kim. 2023 · 2023
Cited alongside, same era.
Towards personalizing generative ai with small data for co-creation in the visual arts
Ahmed M Abuzuraiq and Philippe Pasquier. 2024 · 2024
Cited alongside, same era.
Myvlm: Personalizing vlms for user-specific queries
Yuval Alaluf, Elad Richardson, Sergey Tulyakov, Kfir Aberman, and Daniel Cohen-Or. 2024 · 2024
Cited alongside, same era.
Multimodal large language models in health care: Applications, challenges, and future outlook
Rawan AlSaad, Alaa Abd-Alrazaq, Sabri Boughorbel, Arfan Ahmed, Max-Antoine Renault, Rafat Damseh, and Javaid Sheikh. 2024 · 2024
Cited alongside, same era.
Closest in time.
Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning
Jian Ma, Junhao Liang, Chen Chen, and Haonan Lu. 2024 · 2024
Closest in time.
Yo’llava: Your personalized language and vision assistant
Thao Nguyen, Haotian Liu, Yuheng Li, Mu Cai, Utkarsh Ojha, and Yong Jae Lee. 2024 · 2024
Closest in time.
Maitreya Patel, Sangmin Jung, Chitta Baral, and Yezhou Yang. 2024 · 2024
Closest in time.
Personalized visual instruction tuning
Renjie Pi, Jianshu Zhang, Tianyang Han, Jipeng Zhang, Rui Pan, and Tong Zhang. 2024 · 2024
Closest in time.
Concon-chi: Concept-context chimera benchmark for personalized vision-language tasks
Andrea Rosasco, Stefano Berti, Giulia Pasquale, Damiano Malafronte, Shogo Sato, Hiroyuki Segawa, Tetsugo Inada, and Lorenzo Natale. 2024 · 2024
Closest in time.
Pmg: Personalized multimodal generation with large language models
Xiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu, and Xi Xiao. 2024 · 2024
Closest in time.
Moma: Multimodal llm adapter for fast personalized image generation
Kunpeng Song, Yizhe Zhu, Bingchen Liu, Qing Yan, Ahmed Elgammal, and Xiao Yang. 2024 · 2024
Closest in time.
Tomoya Sugihara, Shuntaro Masuda, Ling Xiao, and Toshihiko Yamasaki. 2024 · 2024
Closest in time.
Mmrec: Llm based multi-modal recommender system
Jiahao Tian, Jinman Zhao, Zhenkai Wang, and Zhicheng Ding. 2024 · 2024
Closest in time.
Multimodal query suggestion with multi-agent reinforcement learning from human feedback
Zheng Wang, Bingzheng Gan, and Wei Shi. 2024e · 2024
Closest in time.
Understanding human preferences: Towards more personalized video to text generation
Yihan Wu, Ruihua Song, Xu Chen, Hao Jiang, Zhao Cao, and Jin Yu. 2024b · 2024
Closest in time.
Jingyi Xie, Rui Yu, He Zhang, Sooyeon Lee, Syed Masum Billah, and John M Carroll. 2024 · 2024
Closest in time.
Align and retrieve: Composition and decomposition learning in image retrieval with text feedback
Yahui Xu, Yi Bin, Jiwei Wei, Yang Yang, Guoqing Wang, and Heng Tao Shen. 2024 · 2024
Closest in time.
Ra-rec: An efficient id representation alignment framework for llm-based recommendation
Xiaohan Yu, Li Zhang, Xin Zhao, Yue Wang, and Zhongrui Ma. 2024 · 2024
Closest in time.
Gpt4rec: Graph prompt tuning for streaming recommendation
Peiyan Zhang, Yuchen Yan, Xi Zhang, Liying Kang, Chaozhuo Li, Feiran Huang, Senzhang Wang, and Sunghun Kim. 2024 · 2024
Closest in time.
User-friendly customized generation with multi-modal prompts
Linhao Zhong, Yan Hong, Wentao Chen, Binglin Zhou, Yiyi Zhang, Jianfu Zhang, and Liqing Zhang. 2024 · 2024
Closest in time.