Fetching the paper…
Reading the bibliography…
Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval.
WordNet: A Lexical Database for English
George A. Miller. 1994 · 1994
Earlier work this paper cites.
Germanet-a lexical-semantic net for german
Birgit Hamp and Helmut Feldweg. 1997 · 1997
Earlier work this paper cites.
Multiwordnet: developing an aligned multilingual database
Emanuele Pianta, Luisa Bentivogli, and Christian Girardi. 2002 · 2002
Earlier work this paper cites.
Automated Flower Classification over a Large Number of Classes
Maria-Elena Nilsback and Andrew Zisserman. 2008 · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
BabelNet: Building a Very Large Multilingual Semantic Network
Roberto Navigli and Simone Paolo Ponzetto. 2010 · 2010
Earlier work this paper cites.
Towards Zero-shot Cross-lingual Image Retrieval
Pranav Aggarwal and Ajinkya Kale. 2020 · 2012
Earlier work this paper cites.
Cats and dogs
Omkar M. Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V. Jawahar. 2012 · 2012
Earlier work this paper cites.
Wikidata: A new platform for collaborative data collection
Denny Vrandečić. 2012 · 2012
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Multi30K: Multilingual English-German Image Descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description
Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Image Pivoting for Learning Multilingual Multimodal Representations
Spandana Gella, Rico Sennrich, Frank Keller, and Mirella Lapata. 2017 · 2017
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
A. Karpathy and L. Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D. Sculley. 2017 · 2017
Earlier work this paper cites.
STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset
Yuya Yoshikawa, Yutaro Shigeto, and Akikazu Takeuchi. 2017 · 2017
Earlier work this paper cites.
Findings of the Third Shared Task on Multimodal Machine Translation
Loïc Barrault, Fethi Bougares, Lucia Specia, Chiraag Lala, Desmond Elliott, and Stella Frank. 2018 · 2018
Earlier work this paper cites.
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Earlier work this paper cites.
Linking ImageNet WordNet Synsets with Wikidata
Finn Årup Nielsen. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Does Object Recognition Work for Everyone?
Terrance DeVries, Ishan Misra, Changhan Wang, and Laurens van der Maaten. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
COCO-CN for Cross-Lingual Image Tagging, Captioning, and Retrieval
Xirong Li, Chaoxi Xu, Xiaoxu Wang, Weiyu Lan, Zhengxiong Jia, Gang Yang, and Jieping Xu. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Do ImageNet Classifiers Generalize to ImageNet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. 2019 · 2019
Cited alongside, same era.
Language-Agnostic Visual-Semantic Embeddings
Jonatas Wehrmann, Maurício Armani Lopes, Douglas M. Souza, and Rodrigo C. Barros. 2019 · 2019
Cited alongside, same era.
Learning to Scale Multilingual Representations for Vision-Language Tasks
Andrea Burns, Donghyun Kim, Derry Wijaya, Kate Saenko, and Bryan A. Plummer. 2020 · 2020
Cited alongside, same era.
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021 · 2021
Later among the works it cites.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2021 · 2021
Later among the works it cites.
WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021 · 2021
Later among the works it cites.
UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-Training
Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, and Jingjing Liu. 2021 · 2021
Later among the works it cites.
IGLUE: A benchmark for transfer learning across modalities, tasks, and languages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
MULE: Multimodal Universal Language Embedding
Donghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff, and Bryan A. Plummer. 2020 · 2020
Cited alongside, same era.
From zero to hero: On the limitations of zero-shot language transfer with multilingual transformers
Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020 · 2020
Cited alongside, same era.
AdapterHub: A Framework for Adapting Transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020a · 2020
Cited alongside, same era.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2020b · 2020
Cited alongside, same era.
Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation
Nils Reimers and Iryna Gurevych. 2020 · 2020
Cited alongside, same era.
Contrastive Language-Image Pre-training for the Italian Language
Federico Bianchi, Giuseppe Attanasio, Raphael Pisoni, Silvia Terragni, Gabriele Sarti, and Sri Lakshmi. 2021 · 2021
Cited alongside, same era.
Emanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, and Ivan Vulic. 2022 · 2022
Later among the works it cites.
Cross-lingual and Multilingual CLIP
Fredrik Carlsson, Philipp Eisen, Faton Rekathati, and Magnus Sahlgren. 2022 · 2022
Later among the works it cites.
No Language Left Behind: Scaling Human-Centered Machine Translation
Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loïc Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Later among the works it cites.
MAGMA - multimodal augmentation of generative models through adapter-based finetuning
Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu, and Anette Frank. 2022 · 2022
Later among the works it cites.
Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval
Gregor Geigle, Jonas Pfeiffer, Nils Reimers, Ivan Vulic, and Iryna Gurevych. 2022 · 2022
Later among the works it cites.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Later among the works it cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022 · 2022
Later among the works it cites.
xGQA: Cross-Lingual Visual Question Answering
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin O. Steitz, Stefan Roth, Ivan Vulic, and Iryna Gurevych. 2022 · 2022
Later among the works it cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022 · 2022
Later among the works it cites.
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
Ashish V. Thapliyal, Jordi Pont-Tuset, Xi Chen, and Radu Soricut. 2022 · 2022
Later among the works it cites.
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
An Yang, Junshu Pan, Junyang Lin, Rui Men, Yichang Zhang, Jingren Zhou, and Chang Zhou. 2022 · 2022
Later among the works it cites.
Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training
Yan Zeng, Wangchunshu Zhou, Ao Luo, and Xinsong Zhang. 2022 · 2022
Later among the works it cites.
LiT: Zero-Shot Transfer with Locked-image text Tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. 2022 · 2022
Later among the works it cites.
Generalizing Multimodal Pre-training into Multilingual via Language Acquisition
Liang Zhang, Anwen Hu, and Qin Jin. 2022 · 2022
Later among the works it cites.
Does progress on ImageNet transfer to real-world datasets?
Alex Fang, Simon Kornblith, and Ludwig Schmidt. 2023 · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023 · 2023
Closest in time.
NLLB-CLIP - train performant multilingual image retrieval model on a budget
Alexander A. Visheratin. 2023 · 2023
Closest in time.
Sigmoid Loss for Language Image Pre-Training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023 · 2023
Closest in time.