Fetching the paper…
Reading the bibliography…
We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2016
Earlier work this paper cites.
Densecap: Fully convolutional localization networks for dense captioning
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Earlier work this paper cites.
Yfcc100m: the new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Visual dialog
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Dense captioning with joint inference and visual context
L. Yang, K. Tang, J. Yang, and L.-J. Li · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
Language-conditioned graph networks for relational reasoning
R. Hu, A. Rohrbach, T. Darrell, and K. Saenko · 2019
Earlier work this paper cites.
Learning by abstraction: The neural state machine
D. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Gqa: a new dataset for compositional question answering over real-world images
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Dense relational captioning: Triple-stream networks for relationship-based captioning
D.-J. Kim, J. Choi, T.-H. Oh, and I. S. Kweon · 2019
Earlier work this paper cites.
Learning object context for dense captioning
X. Li, S. Jiang, and J. Han · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
Context and attribute grounded dense captioning
G. Yin, L. Sheng, B. Liu, N. Yu, X. Wang, and J. Shao · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Length-controllable image captioning
C. Deng, N. Ding, M. Tan, and Q. Wu · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Open-vocabulary object detection via vision and language knowledge distillation
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
An empirical study of gpt-3 for few-shot knowledge-based vqa
Z. Yang, Z. Gan, J. Wang, X. Hu, Y. Lu, Z. Liu, and L. Wang · 2022
Later among the works it cites.
Regionclip: Region-based language-image pretraining
Y. Zhong, J. Yang, P. Zhang, C. Li, N. Codella, L. H. Li, L. Zhou, X. Dai, L. Yuan, Y. Li, et al · 2022
Later among the works it cites.
Getting vit in shape: Scaling laws for compute-optimal model design
I. Alabdulmohsin, X. Zhai, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Palm 2 technical report
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Later among the works it cites.
PaLI: A jointly-scaled multilingual language-image model
X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, A. Kolesnikov, J. Puigcerver, N. Ding, K. Rong, H. Akbari, G. Mishra, L. Xue, A. V. Thapliyal, J. Bradbury, W. Kuo, M. Seyedhosseini, C. Jia, B. K. Ayan, C. R. Ruiz, A. P. Steiner, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
W. Jin, Y. Cheng, Y. Shen, W. Chen, and X. Ren · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Simvlm: Simple visual language model pretraining with weak supervision
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
Flamingo: a visual language model for Few-Shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei · 2022
Cited alongside, same era.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Later among the works it cites.
Promptcap: Prompt-guided image captioning for vqa with gpt-3
Y. Hu, H. Hua, Z. Yang, W. Shi, N. A. Smith, and J. Luo · 2023
Later among the works it cites.
BLIP-2: Bootstrapping Language-Image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Later among the works it cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Later among the works it cites.
Scaling open-vocabulary object detection
M. Minderer, A. Gritsenko, and N. Houlsby · 2023
Later among the works it cites.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Z. Shao, Z. Yu, M. Wang, and J. Yu · 2023
Later among the works it cites.
ViperGPT: Visual inference via python execution for reasoning
D. Sur’is, S. Menon, and C. Vondrick · 2023
Later among the works it cites.
Caption anything: Interactive image description with diverse multimodal controls
T. Wang, J. Zhang, J. Fei, Y. Ge, H. Zheng, Y. Tang, Z. Li, M. Gao, S. Zhao, Y. Shan, et al · 2023
Later among the works it cites.
The all-seeing project: Towards panoptic visual recognition and understanding of the open world
W. Wang, M. Shi, Q. Li, W. Wang, Z. Huang, L. Xing, Z. Chen, H. Li, X. Zhu, Z. Cao, et al · 2023
Later among the works it cites.
Ferret: Refer and ground anything anywhere at any granularity
H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y. Yang · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Instructblip: towards general-purpose vision-language models with instruction tuning
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2024
Closest in time.
Grounding multimodal large language models to the world
Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, Q. Ye, and F. Wei · 2024
Closest in time.
Emu: Generative pretraining in multimodality
Q. Sun, Q. Yu, Y. Cui, F. Zhang, X. Zhang, Y. Wang, H. Gao, J. Liu, T. Huang, and X. Wang · 2024
Closest in time.
Clim: Contrastive language-image mosaic for region representation
S. Wu, W. Zhang, L. Xu, S. Jin, W. Liu, and C. C. Loy · 2024
Closest in time.