Fetching the paper…
Reading the bibliography…
Having revolutionized natural language processing (NLP) applications, large language models (LLMs) are expanding into the realm of multimodal inputs.
“Microsoft coco: Common objects in context,”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, · 2014
Earlier work this paper cites.
“Meteor universal: Language specific translation evaluation for any target language,”
Michael Denkowski and Alon Lavie, · 2014
Earlier work this paper cites.
“Generation and comprehension of unambiguous object descriptions,”
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy, · 2016
Earlier work this paper cites.
“Modeling context in referring expressions,”
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg, · 2016
Earlier work this paper cites.
“Making the v in vqa matter: Elevating the role of image understanding in visual question answering,”
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh, · 2017
Earlier work this paper cites.
“Affectnet: A database for facial expression, valence, and arousal computing in the wild,”
Ali Mollahosseini, Behzad Hasani, and Mohammad H Mahoor, · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“Referring expression generation and comprehension via attributes,”
Jingyu Liu, Liang Wang, and Ming-Hsuan Yang, · 2017
Cited alongside, same era.
“Ok-vqa: A visual question answering benchmark requiring external knowledge,”
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi, · 2019
Cited alongside, same era.
“Gqa: A new dataset for real-world visual reasoning and compositional question answering,”
Drew A Hudson and Christopher D Manning, · 2019
Cited alongside, same era.
“Deep neural network augmentation: Generating faces for affect analysis,”
Dimitrios Kollias, Shiyang Cheng, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou, · 2020
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning,”
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., · 2022
Cited alongside, same era.
“Hagrid-hand gesture recognition image dataset,”
Alexander Kapitanov, Andrew Makhlyarchuk, and Karina Kvanchiani, · 2022
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi, · 2023
Later among the works it cites.
“Instructblip: Towards general-purpose vision-language models with instruction tuning,”
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi, · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dongxu Li, Junnan Li, Hung Le, Guangsen Wang, Silvio Savarese, and Steven CH Hoi, · 2022
Cited alongside, same era.
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, and Ludwig Schmidt, · 2023
Later among the works it cites.
“Kosmos-2: Grounding multimodal large language models to the world,”
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei, · 2023
Later among the works it cites.