Fetching the paper…
Reading the bibliography…
Understanding the mechanisms of information storage and transfer in Transformer-based models is important for driving model understanding progress.
OK-VQA: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 1906
Earlier work this paper cites.
Correlation Matrix Memories
T. Kohonen · 1972
Earlier work this paper cites.
Direct and Indirect Effects
J. Pearl · 2001
Earlier work this paper cites.
Causal Mediation Analysis for Interpreting Neural NLP: The Case of Gender Bias
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. M. Shieber · 2004
Earlier work this paper cites.
Modifying memories in transformer models
C. Zhu, A. S. Rawat, M. Zaheer, S. Bhojanapalli, D. Li, F. X. Yu, and S. Kumar · 2012
Earlier work this paper cites.
VQA: Visual Question Answering, 2016
A. Agrawal, J. Lu, S. Antol, M. Mitchell, C. L. Zitnick, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, and M. Reynolds · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng and David Bau and Alex Andonian and Yonatan Belinkov · 2022
Earlier work this paper cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, 2022
J. Li, D. Li, C. Xiong, and S. Hoi · 2022
Earlier work this paper cites.
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small, 2022
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Cited alongside, same era.
CoCa: Contrastive Captioners are Image-Text Foundation Models
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu · 2022
Cited alongside, same era.
Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization, 2023
A. Agrawal, I. Kajić, E. Bugliarello, E. Davoodi, A. Gergely, P. Blunsom, and A. Nematzadeh · 2023
Cited alongside, same era.
Localizing and Editing Knowledge in Text-to-Image Generative Models, 2023
S. Basu, N. Zhao, V. Morariu, S. Feizi, and V. Manjunatha · 2023
Cited alongside, same era.
Dissecting Recall of Factual Associations in Auto-Regressive Language Models, 2023
M. Geva, J. Bastings, K. Filippova, and A. Globerson · 2023
Cited alongside, same era.
The dawn of LMMs: Preliminary explorations with GPT-4V (ision)
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
M. Yuksekgonul, V. Chandrasekaran, E. Jones, S. Gunasekar, R. Naik, H. Palangi, E. Kamar, and B. Nushi · 2023
Later among the works it cites.
On Mechanistic Knowledge Localization in Text-to-Image Generative Models, 2024
S. Basu, K. Rezaei, P. Kattakinda, R. Rossi, C. Zhao, V. Morariu, V. Manjunatha, and S. Feizi · 2024
Closest in time.
Can We Edit Multimodal Large Language Models?, 2024
S. Cheng, B. Tian, Q. Liu, X. Chen, Y. Wang, H. Chen, and N. Zhang · 2024
Closest in time.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models, 2023
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun · 2023
Cited alongside, same era.
C. Lawless, J. Schoeffer, L. Le, K. Rowan, S. Sen, C. S. Hill, J. Suh, and B. Sarrafzadeh · 2023
Cited alongside, same era.
Improved Baselines with Visual Instruction Tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Cited alongside, same era.
Bounding the capabilities of large language models in open text generation with prompt constraints
A. Lu, H. Zhang, Y. Zhang, X. Wang, and D. Yang · 2023
Cited alongside, same era.
Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
T. Mensink, J. Uijlings, L. Castrejon, A. Goel, F. Cadar, H. Zhou, F. Sha, A. Araujo, and V. Ferrari · 2023
Cited alongside, same era.
Gpt-4v(ision) system card, 9 2023
OpenAI · 2023
Cited alongside, same era.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi
Cited in the paper.
M. He, Y. Liu, B. Wu, J. Yuan, Y. Wang, T. Huang, and B. Zhao · 2024
Closest in time.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Locating and Editing Factual Associations in Mamba
A. S. Sharma, D. Atkinson, and D. Bau · 2024
Closest in time.
LVLM-Intrepret: An Interpretability Tool for Large Vision-Language Models, 2024
G. B. M. Stan, R. Y. Rohekar, Y. Gurwicz, M. L. Olson, A. Bhiwandiwalla, E. Aflalo, C. Wu, N. Duan, S.-Y. Tseng, and V. Lal · 2024
Closest in time.
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs, 2024
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie · 2024
Closest in time.