Fetching the paper…
Reading the bibliography…
Large vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understand image inputs for different queries and conduct subsequent reasoning.
Visualbert: A simple and performant baseline for vision and language
Li, L. H · 1908
Earlier work this paper cites.
Revisiting self-training for neural sequence generation
He, J · 1909
Earlier work this paper cites.
Theoretical analysis of self-training with deep networks on unlabeled data
Wei, C · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y · 2014
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J · 2017
Earlier work this paper cites.
Towards vqa models that can read
Singh, A · 2019
Earlier work this paper cites.
LXMERT: Learning cross-modality encoder representations from transformers
Tan, H · 2019
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Li, X · 2020
Earlier work this paper cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Sohn, K · 2020
Earlier work this paper cites.
Self-training with noisy student improves imagenet classification
Xie, Q · 2020
Earlier work this paper cites.
Rethinking pre-training and self-training
Zoph, B · 2020
Earlier work this paper cites.
Large-margin contrastive learning with distance polarization regularizer
Chen, S · 2021
Earlier work this paper cites.
Exploring simple siamese representation learning
Chen, X · 2021
Earlier work this paper cites.
Multi-task self-training for learning general representations
Ghiasi, G · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C · 2021
Earlier work this paper cites.
ViLT: Vision-and-language transformer without convolution or region supervision
Kim, W · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A · 2021
Earlier work this paper cites.
FLAVA: A foundational language and vision alignment model
Singh, A · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y · 2022
Earlier work this paper cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Bao, H · 2022
Earlier work this paper cites.
Pali: A jointly-scaled multilingual language-image model
Chen, X · 2022
Cited alongside, same era.
Cyclip: Cyclic contrastive language-image pretraining
Goel, S · 2022
Cited alongside, same era.
Grounded language-image pre-training
Li, L. H · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Masry, A · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L · 2022
Cited alongside, same era.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Zhao, Y · 2023
Later among the works it cites.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Ahmadian, A · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How much can clip benefit vision-and-language tasks?
Shen, S · 2022
Cited alongside, same era.
The cringe loss: Learning what language not to model
Adolphs, L · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
Bai, J · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S · 2023
Cited alongside, same era.
Rephrase and respond: Let large language models ask better questions for themselves
Deng, Y · 2023
Cited alongside, same era.
Jiang, A. Q · 2023
Cited alongside, same era.
Chen, Z · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W · 2024
Closest in time.
Kto: Model alignment as prospect theoretic optimization
Ethayarajh, K · 2024
Closest in time.
Fränken, J.-P · 2024
Closest in time.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Gao, P · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge
Liu, H · 2024
Closest in time.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
McKinzie, B · 2024
Closest in time.
Iterative reasoning preference optimization
Pang, R. Y · 2024
Closest in time.
Strengthening multimodal large language model with bootstrapped preference optimization
Pi, R · 2024
Closest in time.
Direct nash optimization: Teaching language models to self-improve with general preferences
Rosset, C · 2024
Closest in time.
Gpt-4v (ision) is a human-aligned evaluator for text-to-3d generation
Wu, T · 2024
Closest in time.
Grok-1.5 vision preview
xAI · 2024
Closest in time.
Self-rewarding language models
Yuan, W · 2024
Closest in time.
Weak-to-strong extrapolation expedites alignment
Zheng, C · 2024
Closest in time.
Aligning modalities in vision large language models via preference fine-tuning
Zhou, Y · 2024
Closest in time.