Fetching the paper…
Reading the bibliography…
The advancement of Multimodal Large Language Models (MLLMs) has greatly accelerated the development of applications in understanding integrated texts and images.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav. 2019 · 1901
Earlier work this paper cites.
A phrase-based alignment model for natural language inference
Bill MacCartney, Michel Galley, and Christopher D. Manning. 2008 · 2008
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Fusing local and global features for high-resolution scene classification
Xiaoyong Bian, Chen Chen, Long Tian, and Qian Du. 2017 · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
Compare, compress and propagate: Enhancing neural architectures with alignment factorization for natural language inference
Yi Tay, Anh Tuan Luu, and Siu Cheung Hui. 2018 · 2018
Earlier work this paper cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
An end-to-end local-global-fusion feature extraction network for remote sensing image scene classification
Yafei Lv, Xiaohan Zhang, Wei Xiong, Yaqi Cui, and Mi Cai. 2019 · 2019
Cited alongside, same era.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
Finding universal grammatical relations in multilingual BERT
Ethan A. Chi, John Hewitt, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Probing multimodal embeddings for linguistic properties: the visual-semantic case
Adam Dahlgren Lindström, Johanna Björklund, Suna Bensch, and Frank Drewes. 2020 · 2020
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022 · 2022
Later among the works it cites.
Probing cross-modal semantics alignment capability from the textual perspective
Zheng Ma, Shi Zong, Mianzhi Pan, Jianbing Zhang, Shujian Huang, Xinyu Dai, and Jiajun Chen. 2022 · 2022
Later among the works it cites.
Fuse local and global semantics in representation learning
Yuchi Zhao and Yuhao Zhou. 2022 · 2022
Later among the works it cites.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Unified language-vision pretraining in llm with dynamic discrete visual tokenization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
oLMpics-On What Language Model Pre-training Captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Cited alongside, same era.
Probing pretrained language models for lexical semantics
Ivan Vulić, Edoardo Maria Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
Rethinking local and global feature representation for semantic segmentation
Mohan Chen, Xinxuan Zhao, Bingfei Fu, Li Zhang, and Xiangyang Xue. 2021 · 2021
Cited alongside, same era.
Discourse probing of pretrained language models
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
A Primer in BERTology: What We Know About How BERT Works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022 · 2022
Cited alongside, same era.
Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Chao Liao, Jianchao Tan, Quzhe Huang, Bin Chen, Chenyi Lei, An Liu, Chengru Song, Xiaoqiang Lei, Di Zhang, Wenwu Ou, Kun Gai, and Yadong Mu. 2023 · 2023
Later among the works it cites.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023 · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. 2023 · 2023
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2023 · 2023
Later among the works it cites.
Generative pretraining in multimodality
Quan Sun, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, Yueze Wang, Hongcheng Gao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
OpenAI. 2024 · 2024
Closest in time.