Fetching the paper…
Reading the bibliography…
In this work, we propose GLOV, which enables Large Language Models (LLMs) to act as implicit optimizers for Vision-Language Models (VLMs) to enhance downstream vision tasks.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
A Mathematical Theory of Communication
Shannon, C. E · 1954
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Learning Generative Visual Models from Few Training Examples: An Incremental Bayesian Approach Tested on 101 Object Categories
Fei-Fei, L., Fergus, R., and Perona, P · 2004
Earlier work this paper cites.
Framewise phoneme classification with bidirectional LSTM and other neural network architectures
Graves, A. and Schmidhuber, J · 2005
Earlier work this paper cites.
Automated Flower Classification Over a Large Number of Classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
SUN Database: Large-scale Scene Recognition from Abbey to Zoo
Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A · 2010
Earlier work this paper cites.
Cats and dogs
Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V · 2012
Earlier work this paper cites.
UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
3D Object Representations for Fine-Grained Categorization
Krause, J., Stark, M., Deng, J., and Fei-Fei, L · 2013
Earlier work this paper cites.
Fine-Grained Visual Classification of Aircraft
Maji, S., Rahtu, E., Kannala, J., Blaschko, M., and Vedaldi, A · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Food-101 – Mining Discriminative Components with Random Forests
Bossard, L., Guillaumin, M., and Van Gool, L · 2014
Earlier work this paper cites.
Describing Textures in the Wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Earlier work this paper cites.
Remote Sensing Image Scene Classification: Benchmark and State of the Art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
Helber, P., Bischke, B., Dengel, A., and Borth, D · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Do ImageNet Classifiers Generalize to ImageNet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Learning Robust Global Representations by Penalizing Local Predictive Power
Wang, H., Ge, S., Lipton, Z., and Xing, E. P · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2020
Cited alongside, same era.
Quantifying the contextualization of word representations with semantic class probing
Zhao, M., Dufter, P., Yaghoobzadeh, Y., and Schütze, H · 2020
Cited alongside, same era.
SEED: Self-supervised Distillation for Visual Representation
Fang, Z., Wang, J., Wang, L., Zhang, L., Yang, Y., and Liu, Z · 2021
Cited alongside, same era.
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T · 2021
Cited alongside, same era.
What does a platypus look like? Generating customized prompts for zero-shot image classification
Pratt, S., Liu, R., and Farhadi, A · 2023
Later among the works it cites.
Waffling around for Performance: Visual Classification with Random Words and Broad Concepts
Roth, K., Kim, J. M., Koepke, A., Vinyals, O., Schmid, C., and Akata, Z · 2023
Later among the works it cites.
Dept: Decomposed prompt tuning for parameter-efficient fine-tuning, 2023
Shi, Z. and Lipani, A · 2023
Later among the works it cites.
Function vectors in large language models
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Turner, A., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Power of Scale for Parameter-efficient Prompt Tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X., Liang, P., and Jurafsky, D · 2021
Cited alongside, same era.
Learning Transferable Visual Models from Natural Language Supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Factual Probing is [MASK]: Learning vs. Learning to Recall
Zhong, Z., Friedman, D., and Chen, D · 2021
Cited alongside, same era.
Flamingo: a Visual Language Model for Few-Shot Learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Unsupervised Prompt Learning for Vision-Language Models
Huang, T., Chu, J., and Wei, F · 2022
Cited alongside, same era.
Visual Prompt Tuning
Jia, M., Tang, L., Chen, B.-C., Cardie, C., Belongie, S., Hariharan, B., and Lim, S.-N · 2022
Cited alongside, same era.
Later among the works it cites.
Demystifying CLIP Data
Xu, H., Xie, S., Tan, X. E., Huang, P.-Y., Howes, R., Sharma, V., Li, S.-W., Ghosh, G., Zettlemoyer, L., and Feichtenhofer, C · 2023
Later among the works it cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2023
Later among the works it cites.
A survey on multimodal large language models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2023
Later among the works it cites.
Sigmoid Loss for Language Image Pre-training
Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L · 2023
Later among the works it cites.
xLSTM: Extended Long Short-Term Memory
Beck, M., Pöppel, K., Spanring, M., Auer, A., Prudnikova, O., Kopp, M., Klambauer, G., Brandstetter, J., and Hochreiter, S · 2024
Closest in time.
MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-task Learning
Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., Krishnamoorthi, R., Chandra, V., Xiong, Y., and Elhoseiny, M · 2024
Closest in time.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
Deitke, M., Clark, C., Lee, S., Tripathi, R., Yang, Y., Park, J. S., Salehi, M., Muennighoff, N., Lo, K., Soldaini, L., et al · 2024
Closest in time.
Towards multimodal in-context learning for vision & language models
Doveh, S., Perek, S., Mirza, M. J., Alfassy, A., Arbelle, A., Ullman, S., and Karlinsky, L · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
CLIP-Adapter: Better Vision-language Models with Feature Adapters
Gao, P., Geng, S., Zhang, R., Ma, T., Fang, R., Zhang, Y., Li, H., and Qiao, Y · 2024
Closest in time.
Are Vision Language Models Texture or Shape Biased and Can We Steer Them?
Gavrikov, P., Lukasik, J., Jung, S., Geirhos, R., Lamm, B., Mirza, M. J., Keuper, M., and Keuper, J · 2024
Closest in time.
Geigle, G., Timofte, R., and Glavaš, G · 2024
Closest in time.
Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation
Gou, Y., Chen, K., Liu, Z., Hong, L., Xu, H., Li, Z., Yeung, D.-Y., Kwok, J. T., and Zhang, Y · 2024
Closest in time.
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
Huang, I., Lin, W., Mirza, M. J., Hansen, J. A., Doveh, S., Butoi, V. I., Herzig, R., Arbelle, A., Kuhene, H., Darrel, T., et al · 2024
Closest in time.
Prompt tuning strikes back: Customizing foundation models with low-rank prompt adaptation, 2024
Jain, A., Chaudhuri, S., Reps, T., and Jermaine, C · 2024
Closest in time.
Aapl: Adding attributes to prompt learning for vision-language models
Kim, G., Kim, S., and Lee, S · 2024
Closest in time.
LLaVA-OneVision: Easy Visual Task Transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., and Li, C · 2024
Closest in time.
Comparison Visual Instruction Tuning
Lin, W., Mirza, M. J., Doveh, S., Feris, R., Giryes, R., Hochreiter, S., and Karlinsky, L · 2024
Closest in time.
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
Mirza, M. J., Karlinsky, L., Lin, W., Doveh, S., , Micorek, J., Kozinski, M., Kuhene, H., and Possegger, H · 2024
Closest in time.
How Many Unicorns Are In This Image? A Safety Evaluation Benchmark For Vision LLMs
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C · 2024
Closest in time.
Large Language Models as Optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2024
Closest in time.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2024
Closest in time.
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
Zong, Y., Bohdal, O., Yu, T., Yang, Y., and Timothy, H · 2024
Closest in time.