Fetching the paper…
Reading the bibliography…
Unlike traditional vision-only models, vision language models (VLMs) offer an intuitive way to access visual content through language prompting by combining a large language model (LLM) with a vision encoder.
A Coefficient of Agreement for Nominal Scales
Jacob Cohen · 1960
Earlier work this paper cites.
Hybrid images
Aude Oliva, Antonio Torralba, and Philippe G. Schyns · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Highway Networks, 2015
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Show and Tell: A Neural Image Caption Generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Image Style Transfer Using Convolutional Neural Networks
Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge · 2016
Earlier work this paper cites.
On Calibration of Modern Neural Networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification
Joy Buolamwini and Timnit Gebru · 2018
Earlier work this paper cites.
ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2019
Earlier work this paper cites.
Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products
Inioluwa Deborah Raji and Joy Buolamwini · 2019
Earlier work this paper cites.
PyTorch Image Models
Ross Wightman · 2019
Earlier work this paper cites.
Interpreting Adversarially Trained Convolutional Neural Networks
Tianyuan Zhang and Zhanxing Zhu · 2019
Earlier work this paper cites.
The Origins and Prevalence of Texture Bias in Convolutional Neural Networks
Katherine Hermann, Ting Chen, and Simon Kornblith · 2020
Earlier work this paper cites.
Informative Dropout for Robust Representation Learning: A Shape-bias Perspective
Baifeng Shi, Dinghuai Zhang, Qi Dai, Zhanxing Zhu, Yadong Mu, and Jingdong Wang · 2020
Earlier work this paper cites.
High-Frequency Component Helps Explain the Generalization of Convolutional Neural Networks
Haohan Wang, Xindi Wu, Zeyi Huang, and Eric P. Xing · 2020
Earlier work this paper cites.
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models
Soumya Barikeri, Anne Lauscher, Ivan Vulic, and Goran Glavas · 2021
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Partial success in closing the gap between human and machine vision
Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2021
Earlier work this paper cites.
Multimodal Neurons in Artificial Neural Networks
Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Earlier work this paper cites.
Shape or Texture: Understanding Discriminative Features in CNNs
Md Amirul Islam, Matthew Kowal, Patrick Esser, Sen Jia, Björn Ommer, Konstantinos G. Derpanis, and Neil Bruce · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Cited alongside, same era.
Sustainable Modular Debiasing of Language Models
Anne Lauscher, Tobias Lüken, and Goran Glavas · 2021
Cited alongside, same era.
Shape-Texture Debiased Neural Network Training
Yingwei Li, Qihang Yu, Mingxing Tan, Jieru Mei, Peng Tang, Wei Shen, Alan Yuille, and cihang xie · 2021
Cited alongside, same era.
Learning to Compose Visual Relations
Nan Liu, Shuang Li, Yilun Du, Josh Tenenbaum, and Antonio Torralba · 2021
Cited alongside, same era.
Intriguing Properties of Vision Transformers
Muhammad Muzammal Naseer, Kanchana Ranasinghe, Salman H Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang · 2021
Language Is Not All You Need: Aligning Perception with Language Models
Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei · 2023
Later among the works it cites.
UForm by Unum Cloud, 2023
Mikhail Kim, Vladimir Orshulevich, and Ash Vardanian · 2023
Later among the works it cites.
Improving Native CNN Robustness with Filter Frequency Regularization
Jovita Lukasik, Paul Gavrikov, Janis Keuper, and Margret Keuper · 2023
Later among the works it cites.
Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen, Xiaoshuai Sun, and Rongrong Ji · 2023
Later among the works it cites.
MTEB: Massive Text Embedding Benchmark
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers · 2023
Later among the works it cites.
GPT-4 Technical Report, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Are Convolutional Neural Networks or Transformers more like human vision?, 2021
Shikhar Tuli, Ishita Dasgupta, Erin Grant, and Thomas L. Griffiths · 2021
Cited alongside, same era.
VideoGPT: Video Generation using VQ-VAE and Transformers, 2021
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikoł aj Bińkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karén Simonyan · 2022
Cited alongside, same era.
Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts
Yue Guo, Yi Yang, and Ahmed Abbasi · 2022
Cited alongside, same era.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Evaluating the Moral Beliefs Encoded in LLMs
Nino Scherrer, Claudia Shi, Amir Feder, and David Blei · 2023
Later among the works it cites.
CogVLM: Visual Expert for Pretrained Language Models, 2023
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang · 2023
Later among the works it cites.
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks, 2024
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai · 2024
Closest in time.
shape, 2024a
Dictionary.com · 2024
Closest in time.
texture, 2024b
Dictionary.com · 2024
Closest in time.
Can Biases in ImageNet Models Explain Generalization?
Paul Gavrikov and Janis Keuper · 2024
Closest in time.
Intriguing Properties of Generative Classifiers
Priyank Jaini, Kevin Clark, and Robert Geirhos · 2024
Closest in time.
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models, 2024
Bin Lin, Zhenyu Tang, Yang Ye, Jiaxi Cui, Bin Zhu, Peng Jin, Jinfa Huang, Junwu Zhang, Munan Ning, and Li Yuan · 2024
Closest in time.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
SFR-Embedded-Mistral
Rui Meng, Ye Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz · 2024
Closest in time.
llmrails/ember-v1 ⋅ \cdot Hugging Face, 2024
Enrike Nur and Anar Aliyev · 2024
Closest in time.
Introducing Qwen-VL
Qwen Team · 2024
Closest in time.
Exploring Value Biases: How LLMs Deviate Towards the Ideal, 2024
Sarath Sivaprasad, Pramod Kaushik, Sahar Abdelnabi, and Mario Fritz · 2024
Closest in time.
Spatial-frequency channels, shape bias, and adversarial robustness
Ajay Subramanian, Elena Sizikova, Najib Majaj, and Denis Pelli · 2024
Closest in time.
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters, 2024
Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, and Xinlong Wang · 2024
Closest in time.
Nous Hermes 2 Mistral 7B DPO, 2024
Teknium, theemozilla, karan4d, and huemin_art · 2024
Closest in time.
Large Language Models as Optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen · 2024
Closest in time.