Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs) have gained community-spanning prominence due to their ability to integrate visual and textual inputs to perform complex tasks.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 1908
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 1908
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov. 2020 · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. 2022 · 2010
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. 2011 · 2011
Earlier work this paper cites.
Seeing past words: Testing the cross-modal capabilities of pretrained v&l models on counting tasks
Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto. 2020 · 2012
Earlier work this paper cites.
Challenges in representation learning: A report on three machine learning contests
Ian J Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron Courville, Mehdi Mirza, Ben Hamner, Will Cukierski, Yichuan Tang, David Thaler, Dong-Hyun Lee, et al. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Discovering states and transformations in image collections
Phillip Isola, Joseph J Lim, and Edward H Adelson. 2015 · 2015
Earlier work this paper cites.
Sentiment of emojis
Petra Kralj Novak, Jasmina Smailović, Borut Sluban, and Igor Mozetič. 2015 · 2015
Earlier work this paper cites.
Large-scale fashion (deepfashion) database
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. 2016 · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Earlier work this paper cites.
Shapeworld-a new test methodology for multimodal language understanding
Alexander Kuhnle and Ann Copestake. 2017 · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Earlier work this paper cites.
What does bert learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Earlier work this paper cites.
Behind the scene: Revealing the secrets of pre-trained vision-and-language models
Jize Cao, Zhe Gan, Yu Cheng, Licheng Yu, Yen-Chun Chen, and Jingjing Liu. 2020 · 2020
Earlier work this paper cites.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Jack Hessel and Lillian Lee. 2020 · 2020
Cited alongside, same era.
Lsun-stanford car dataset: enhancing large-scale car image datasets using deep learning for usage in gan training
Tin Kramberger and Božidar Potočnik. 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020 · 2020
Cited alongside, same era.
The explanation game: Explaining machine learning models using shapley values
Luke Merrick and Ankur Taly. 2020 · 2020
Cited alongside, same era.
Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott. 2021 · 2021
Cited alongside, same era.
A toy model of universality: Reverse engineering how networks learn group operations
Bilal Chughtai, Lawrence Chan, and Neel Nanda. 2023 · 2023
Later among the works it cites.
Successor heads: Recurring, interpretable attention heads in the wild
Rhys Gould, Euan Ong, George Ogden, and Arthur Conmy. 2023 · 2023
Later among the works it cites.
Llms as visual explainers: Advancing image classification with evolving visual descriptions
Songhao Han, Le Zhuo, Yue Liao, and Si Liu. 2023 · 2023
Later among the works it cites.
Michael Hanna, Ollie Liu, and Alexandre Variengien. 2023 · 2023
Later among the works it cites.
Multiviz: Towards visualizing and understanding multimodal models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mario Giulianelli, Jacqueline Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2021 · 2021
Cited alongside, same era.
Probing image-language transformers for verb understanding
Lisa Anne Hendricks and Aida Nematzadeh. 2021 · 2021
Cited alongside, same era.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
Pan Lu, Liang Qiu, Jiaqi Chen, Tony Xia, Yizhou Zhao, Wei Zhang, Zhou Yu, Xiaodan Liang, and Song-Chun Zhu. 2021 · 2021
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Automatic captioning for medical imaging (mic): a rapid review of literature
Beddiar Djamila-Romaissa, Oussalah Mourad, and Seppänen Tapio. 2022 · 2022
Cited alongside, same era.
Attention as grounding: Exploring textual and cross-modal attention on entities and relations in language-and-vision transformer
Nikolai Ilinykh and Simon Dobnik. 2022a · 2022
Cited alongside, same era.
Paul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain, Zihao Deng, Xingbo Wang, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2023 · 2023
Later among the works it cites.
Vim: Probing multimodal large language models for visual embedded instruction following
Yujie Lu, Xiujun Li, William Yang Wang, and Yejin Choi. 2023 · 2023
Later among the works it cites.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2023 · 2023
Later among the works it cites.
Towards vision-language mechanistic interpretability: A causal tracing tool for blip
Vedant Palit, Rohan Pandey, Aryaman Arora, and Paul Pu Liang. 2023 · 2023
Later among the works it cites.
Outlier dimensions encode task specific knowledge
William Rudman, Catherine Chen, and Carsten Eickhoff. 2023 · 2023
Later among the works it cites.
Multimodal learning with transformers: A survey
Peng Xu, Xiatian Zhu, and David A. Clifton. 2023 · 2023
Later among the works it cites.
Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023 · 2023
Later among the works it cites.
Impossibility theorems for feature attribution
Blair Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim. 2024 · 2024
Closest in time.
Vision transformers need registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. 2024 · 2024
Closest in time.
Interpreting and editing vision-language representations to mitigate hallucinations
Nick Jiang, Anish Kachinthaya, Suzie Petryk, and Yossi Gandelsman. 2024 · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024 · 2024
Closest in time.
Grace Luo, Trevor Darrell, and Amir Bar. 2024 · 2024
Closest in time.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2024 · 2024
Closest in time.
Towards interpreting visual information processing in vision-language models
Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, and Fazl Barez. 2024 · 2024
Closest in time.
Analyzing vision transformers for image classification in class embedding space
Martina G Vilas, Timothy Schaumlöffel, and Gemma Roig. 2024 · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda. 2024 · 2024
Closest in time.