Fetching the paper…
Reading the bibliography…
PaliGemma 2 is an upgrade of the PaliGemma open Vision-Language Model (VLM) based on the Gemma 2 family of language models.
SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules
D. Weininger · 1988
Earlier work this paper cites.
PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals
A. L. Goldberger, L. A. Amaral, L. Glass, J. M. Hausdorff, P. C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley · 2000
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
D. L. Chen and W. B. Dolan · 2011
Earlier work this paper cites.
Indigo: Universal cheminformatics API
D. Pavlov, M. Rybalkin, B. Karulin, M. Kozhevnikov, A. Savelyev, and A. Churinov · 2011
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
T. Lin, M. Maire, S. J. Belongie, L. D. Bourdev, R. B. Girshick, J. Hays, P. Perona, D. Ramanan, P. Doll’a r, and C. L. Zitnick · 2014
Earlier work this paper cites.
ICDAR 2015 competition on robust reading
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. K. Ghosh, A. D. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny · 2015
Earlier work this paper cites.
Neural combinatorial optimization with reinforcement learning
I. Bello, H. Pham, Q. V. Le, M. Norouzi, and S. Bengio · 2016
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Pubchem substance and compound databases
S. Kim, P. A. Thiessen, E. E. Bolton, J. Chen, G. Fu, A. Gindulyte, L. Han, J. He, S. He, B. A. Shoemaker, et al · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Total-Text: A comprehensive dataset for scene text detection and recognition
C. K. Ch’ng and C. S. Chan · 2017
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles · 2017
Earlier work this paper cites.
ICDAR2017 robust reading challenge on multi-lingual scene text detection and script identification - RRC-MLT
N. Nayef, F. Yin, I. Bizid, H. Choi, Y. Feng, D. Karatzas, Z. Luo, U. Pal, C. Rigaud, J. Chazalon, et al · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
VizWiz Grand Challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
TallyQA: Answering complex counting questions
M. Acharya, K. Kafle, and C. Kanan · 2019
Earlier work this paper cites.
NoCaps: Novel object captioning at scale
H. Agrawal, K. Desai, Y. Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson · 2019
Earlier work this paper cites.
Scene text visual question answering
A. F. Biten, R. Tito, A. Mafla, L. Gomez, M. Rusinol, C. Jawahar, E. Valveny, and D. Karatzas · 2019
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
D. Hudson and C. Manning · 2019
Earlier work this paper cites.
MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports
A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C.-Y. Deng, R. G. Mark, and S. Horng · 2019
Earlier work this paper cites.
OK-VQA: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
OCR-VQA: Visual question answering by reading text in images
A. Mishra, S. Shekhar, A. K. Singh, and A. Chakraborty · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
A corpus for reasoning about natural language grounded in photographs
A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi · 2019
Earlier work this paper cites.
LXMERT: Learning cross-modality encoder representations from transformers
H. Tan and M. Bansal · 2019
Earlier work this paper cites.
VaTeX: A large-scale, high-quality multilingual dataset for video-and-language research
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang · 2019
Earlier work this paper cites.
ActivityNet-QA: A dataset for understanding complex web videos via question answering
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao · 2019
Earlier work this paper cites.
Widget Captioning: Generating natural language description for mobileuser interface elements
Y. Li, G. Li, L. He, J. Zheng, H. Li, and Z. Guan · 2020
Earlier work this paper cites.
RSVQA: Visual question answering for remote sensing data
S. Lobry, D. Marcos, J. Murray, and D. Tuia · 2020
Earlier work this paper cites.
DocVQA: A dataset for VQA on document images
M. Mathew, D. Karatzas, R. Manmatha, and C. V. Jawahar · 2020
Earlier work this paper cites.
TextCaps: A dataset for image captioning with reading comprehension
O. Sidorov, R. Hu, M. Rohrbach, and A. Singh · 2020
Cited alongside, same era.
Image-based table recognition: Data, model, and evaluation
X. Zhong, E. ShafieiBavani, and A. Jimeno Yepes · 2020
Cited alongside, same era.
Virtex: Learning visual representations from textual annotations
K. Desai and J. Johnson · 2021
Cited alongside, same era.
Scicap: Generating captions for scientific figures
T.-Y. Hsu, C. L. Giles, and T.-H. Huang · 2021
Cited alongside, same era.
Scaling up visual and vision-language representation learning with noisy text supervision
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig · 2021
Cited alongside, same era.
Ultralytics YOLO, 2023
G. Jocher, J. Qiu, and A. Chaurasia · 2023
Later among the works it cites.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. C. H. Hoi · 2023
Later among the works it cites.
ICDAR 2023 competition on hierarchical text detection and recognition
S. Long, S. Qin, D. Panteleev, A. Bissacco, Y. Fujii, and M. Raptis · 2023
Later among the works it cites.
An end-to-end multi-task learning model for image-based table recognition
N. T. Ly and A. Takasu · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, and F. Wei · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Karkkainen and J. Joo · 2021
Cited alongside, same era.
Open images v5 text annotation and yet another mask text spotter
I. Krylov, S. Nosov, and V. Sovrasov · 2021
Cited alongside, same era.
Visually grounded reasoning across languages and cultures
F. Liu, E. Bugliarello, E. M. Ponti, S. Reddy, N. Collier, and D. Elliott · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner · 2021
Cited alongside, same era.
Screen2words: Automatic mobile ui summarization with multimodal learning
B. Wang, G. Li, X. Zhou, Z. Chen, T. Grossman, and Y. Li · 2021
Cited alongside, same era.
Global Table Extractor (GTE): A framework for joint table identification and cell structure recognition using visual context
X. Zheng, D. Burdick, L. Popa, P. Zhong, and N. X. R. Wang · 2021
Cited alongside, same era.
MolScribe: Robust molecular structure recognition with image-to-graph generation
Y. Qian, J. Guo, Z. Tu, Z. Li, C. W. Coley, and R. Barzilay · 2023
Later among the works it cites.
Measuring attribution in natural language generation models
H. Rashkin, V. Nikolaev, M. Lamm, L. Aroyo, M. Collins, D. Das, S. Petrov, G. S. Tomar, I. Turc, and D. Reitter · 2023
Later among the works it cites.
End-to-end optical music recognition for pianoform sheet music
A. Ríos-Vila, D. Rizo, J. M. Iñesta, and J. Calvo-Zaragoza · 2023
Later among the works it cites.
Aligning benchmark datasets for table structure recognition
B. Smock, R. Pesala, and R. Abraham · 2023
Later among the works it cites.
Tuning computer vision models with task rewards
A. Susano Pinto, A. Kolesnikov, Y. Shi, L. Beyer, and X. Zhai · 2023
Later among the works it cites.
Image captioners are scalable vision learners too
M. Tschannen, M. Kumar, A. Steiner, X. Zhai, N. Houlsby, and L. Beyer · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Pytorch FSDP: experiences on scaling fully sharded data parallel
Y. Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li · 2023
Later among the works it cites.
PaliGemma: A versatile 3B VLM for transfer
L. Beyer, A. Steiner, A. S. Pinto, A. Kolesnikov, X. Wang, D. Salz, M. Neumann, I. Alabdulmohsin, M. Tschannen, E. Bugliarello, T. Unterthiner, D. Keysers, S. Koppula, F. Liu, A. Grycner, A. Gritsenko, N. Houlsby, M. Kumar, K. Rong, J. Eisenschlos, R. Kabra, M. Bauer, M. Bošnjak, X. Chen, M. Minderer, P. Voigtlaender, I. Bica, I. Balazevic, J. Puigcerver, P. Papalampidi, O. Henaff, X. Xiong, R. Soricut, J. Harmsen, and X. Zhai · 2024
Closest in time.
PaLI-X: On scaling up a multilingual vision and language model
X. Chen, J. Djolonga, P. Padlewski, B. Mustafa, S. Changpinyo, J. Wu, C. R. Ruiz, S. Goodman, X. Wang, Y. Tay, S. Shakeri, M. Dehghani, D. Salz, M. Lucic, M. Tschannen, A. Nagrani, H. Hu, M. Joshi, B. Pang, C. Montgomery, P. Pietrzyk, M. Ritter, A. J. Piergiovanni, M. Minderer, F. Pavetic, A. Waters, G. Li, I. Alabdulmohsin, L. Beyer, J. Amelot, K. Lee, A. P. Steiner, Y. Li, D. Keysers, A. Arnab, Y. Xu, K. Rong, A. Kolesnikov, M. Seyedhosseini, A. Angelova, X. Zhai, N. Houlsby, and R. Soricut · 2024
Closest in time.
Molmo and PixMo: Open weights and open data for state-of-the-art multimodal models
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al · 2024
Closest in time.
Introduction to Cloud TPU
Google Cloud · 2024
Closest in time.
BRAVE: Broadening the visual encoding of vision-language models
O. F. Kar, A. Tonioni, P. Poklukar, A. Kulshrestha, A. Zamir, and F. Tombari · 2024
Closest in time.
Prismatic VLMs: Investigating the design space of visually-conditioned language models
S. Karamcheti, S. Nair, A. Balakrishna, P. Liang, T. Kollar, and D. Sadigh · 2024
Closest in time.
Multi-cell decoder and mutual learning for table structure and character recognition
T. Kawakatsu · 2024
Closest in time.
What matters when building vision-language models?
H. Laurençon, L. Tronchon, M. Cord, and V. Sanh · 2024
Closest in time.
LLaVA-NeXT: What else influences visual instruction tuning beyond data?, May 2024
B. Li, H. Zhang, K. Zhang, D. Guo, Y. Zhang, R. Zhang, F. Li, Z. Liu, and C. Li · 2024
Closest in time.
Hierarchical text spotter for joint text spotting and layout analysis
S. Long, S. Qin, Y. Fujii, A. Bissacco, and M. Raptis · 2024
Closest in time.
MM1: methods, analysis & insights from multimodal LLM pre-training
B. McKinzie, Z. Gan, J. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, A. Belyi, H. Zhang, K. Singh, D. Kang, A. Jain, H. Hè, M. Schwarzer, T. Gunter, X. Kong, A. Zhang, J. Wang, C. Wang, N. Du, T. Lei, S. Wiseman, G. Yin, M. Lee, Z. Wang, R. Pang, P. Grasch, A. Toshev, and Y. Yang · 2024
Closest in time.
DOCCI: Descriptions of Connected and Contrasting Images
Y. Onoe, S. Rane, Z. Berger, Y. Bitton, J. Cho, R. Garg, A. Ku, Z. Parekh, J. Pont-Tuset, G. Tanzer, S. Wang, and J. Baldridge · 2024
Closest in time.
YOLO-DocLayNet, Jan. 2024
H. Pang · 2024
Closest in time.
Sheet Music Transformer: End-to-end optical music recognition beyond monophonic transcription
A. Ríos-Vila, J. Calvo-Zaragoza, and T. Paquet · 2024
Closest in time.
Collaboration between clinicians and vision–language models in radiology report generation
R. Tanno, D. Barrett, A. Sellergren, S. Ghaisas, S. Dathathri, A. See, J. Welbl, K. Singhal, S. Azizi, T. Tu, M. Schaekermann, R. May, R. Lee, S. Man, Z. Ahmed, S. Mahdavi, Y. Matias, J. Barral, A. Eslami, D. Belgrave, V. Natarajan, S. Shetty, P. Kohli, P.-S. Huang, A. Karthikesalingam, and I. Ktena · 2024
Closest in time.
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
S. Tong, E. Brown, P. Wu, S. Woo, M. Middepogu, S. C. Akula, J. Yang, S. Yang, A. Iyer, X. Pan, A. Wang, R. Fergus, Y. LeCun, and S. Xie · 2024
Closest in time.
LocCa: Visual pretraining with location-aware captioners
B. Wan, M. Tschannen, Y. Xian, F. Pavetic, I. Alabdulmohsin, X. Wang, A. S. Pinto, A. Steiner, L. Beyer, and X. Zhai · 2024
Closest in time.
Advancing multimodal medical capabilities of Gemini
L. Yang, S. Xu, A. Sellergren, T. Kohlberger, Y. Zhou, I. Ktena, A. Kiraly, F. Ahmed, F. Hormozdiari, T. Jaroensri, E. Wang, E. Wulczyn, F. Jamil, T. Guidroz, C. Lau, S. Qiao, Y. Liu, A. Goel, K. Park, A. Agharwal, N. George, Y. Wang, R. Tanno, D. G. T. Barrett, W.-H. Weng, S. S. Mahdavi, K. Saab, T. Tu, S. R. Kalidindi, M. Etemadi, J. Cuadros, G. Sorensen, Y. Matias, K. Chou, G. Corrado, J. Barral, S. Shetty, D. Fleet, S. M. A. Eslami, D. Tse, S. Prabhakara, C. McLean, D. Steiner, R. Pilgrim, C. Kelly, S. Azizi, and D. Golden · 2024
Closest in time.
mPLUG-Owl2: Revolutionizing multi-modal large language model with modality collaboration
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang · 2024
Closest in time.
Ferret: Refer and ground anything anywhere at any granularity
H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y. Yang · 2024
Closest in time.
MM1.5: Methods, analysis & insights from multimodal LLM fine-tuning
H. Zhang, M. Gao, Z. Gan, P. Dufter, N. Wenzel, F. Huang, D. Shah, X. Du, B. Zhang, Y. Li, et al · 2024
Closest in time.