Fetching the paper…
Reading the bibliography…
We investigate the internal representations of vision-language models (VLMs) to address hallucinations, a persistent challenge despite advances in model size and training.
A disentangling invertible interpretation network for explaining latent representations, 2020
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2004
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models, 2018
Vitali Petsiuk, Abir Das, and Kate Saenko · 2018
Earlier work this paper cites.
Gan dissection: Visualizing and understanding generative adversarial networks
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.
Object hallucination in image captioning, 2019
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko · 2019
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui · 2020
Earlier work this paper cites.
Interpreting GPT: The logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Hila Chefer, Shir Gur, and Lior Wolf · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
What if this modified that? syntactic interventions via counterfactual embeddings, 2021
Mycal Tucker, Peng Qian, and Roger Levy · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers, 2022
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Earlier work this paper cites.
The internal state of an LLM knows when it’s lying
Amos Azaria and Tom Mitchell · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Earlier work this paper cites.
Eliciting latent predictions from transformers with the tuned lens, 2023
Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Interpreting and controlling vision foundation models via text explanations, 2023
Haozhe Chen, Junfeng Yang, Carl Vondrick, and Chengzhi Mao · 2023
Cited alongside, same era.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models, 2023
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey · 2023
Cited alongside, same era.
Inside: Llms’ internal states retain the power of hallucination detection, 2024
Chao Chen, Kai Liu, Ze Chen, Yi Gu, Yue Wu, Mingyuan Tao, Zhihang Fu, and Jieping Ye · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models, 2024
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva · 2024
Closest in time.
Overthinking the truth: Understanding how language models process false demonstrations, 2024
Danny Halawi, Jean-Stanislas Denain, and Jacob Steinhardt · 2024
Closest in time.
Llm factoscope: Uncovering llms’ factual discernment through inner states analysis, 2024
Jinwen He, Yujia Gong, Kai Chen, Zijin Lin, Chengan Wei, and Yue Zhao · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Cited alongside, same era.
Rosetta neurons: Mining the common units in a model zoo
Amil Dravid, Yossi Gandelsman, Alexei A. Efros, and Assaf Shocher · 2023
Cited alongside, same era.
Rsgpt: A remote sensing vision language model and benchmark, 2023
Yuan Hu, Jianlong Yuan, Congcong Wen, Xiaonan Lu, and Xiang Li · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Cited alongside, same era.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Cited alongside, same era.
Towards vision-language mechanistic interpretability: A causal tracing tool for blip, 2023
Vedant Palit, Rohan Pandey, Aryaman Arora, and Paul Pu Liang · 2023
Cited alongside, same era.
Multimodal neurons in pretrained text-only transformers, 2023
Sarah Schwettmann, Neil Chowdhury, Samuel Klein, and Antonio Torralba · 2023
Cited alongside, same era.
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Cited alongside, same era.
Qidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang, Conghui He, Jiaqi Wang, Dahua Lin, Weiming Zhang, and Nenghai Yu · 2024
Closest in time.
Jiahao Huo, Yibo Yan, Boren Hu, Yutao Yue, and Xuming Hu · 2024
Closest in time.
Emergent world representations: Exploring a sequence model trained on a synthetic task, 2024
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
Fuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li, Xiaolong Wang, Siyu Wang, Ziyue Wang, Xiaoyue Mi, Peng Li, Ning Ma, Maosong Sun, and Yang Liu · 2024
Closest in time.
Linear adversarial concept erasure, 2024
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan Cotterell · 2024
Closest in time.
Luna: A model-based universal analysis framework for large language models, 2024
Da Song, Xuan Xie, Jiayang Song, Derui Zhu, Yuheng Huang, Felix Juefei-Xu, and Lei Ma · 2024
Closest in time.
Weihang Su, Changyue Wang, Qingyao Ai, Yiran HU, Zhijing Wu, Yujia Zhou, and Yiqun Liu · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie · 2024
Closest in time.
A language model’s guide through latent space, 2024
Dimitri von Rutte, Sotiris Anagnostidis, Gregor Bachmann, and Thomas Hofmann · 2024
Closest in time.
Analyzing and mitigating object hallucination in large vision-language models, 2024
Yiyang Zhou, Chenhang Cui, Jaehong Yoon, Linjun Zhang, Zhun Deng, Chelsea Finn, Mohit Bansal, and Huaxiu Yao · 2024
Closest in time.