Fetching the paper…
Reading the bibliography…
This paper studies zero-shot cross-lingual transfer of vision-language models.
Beto, bentz, becas: The surprising cross-lingual effectiveness of bert
Shijie Wu and Mark Dredze. 2019a · 1904
Earlier work this paper cites.
Unicoder: A universal language encoder by pre-training with multiple cross-lingual tasks
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou. 2019a · 1909
Earlier work this paper cites.
Noise estimation using density estimation for self-supervised multimodal learning
Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein. 2020 · 2003
Earlier work this paper cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2003
Earlier work this paper cites.
Multi-modal self-supervision from generalized data transformations
Mandela Patrick, Y. Asano, Ruth Fong, João F. Henriques, G. Zweig, and A. Vedaldi. 2020 · 2003
Earlier work this paper cites.
Leveraging monolingual data with self-supervision for multilingual neural machine translation
Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, and Yonghui Wu. 2020 · 2005
Earlier work this paper cites.
Video understanding as machine translation
Bruno Korbar, F. Petroni, Rohit Girdhar, and L. Torresani. 2020 · 2006
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
David Chen and William Dolan. 2011 · 2011
Earlier work this paper cites.
Inducing crosslingual distributed representations of words
Alexandre Klementiev, Ivan Titov, and Binod Bhattarai. 2012 · 2012
Earlier work this paper cites.
Globetrotter: Unsupervised multilingual translation from visual alignment
Dídac Surís, Dave Epstein, and Carl Vondrick. 2020 · 2012
Earlier work this paper cites.
Cross-lingual word clusters for direct transfer of linguistic structure
Oscar Täckström, Ryan McDonald, and Jakob Uszkoreit. 2012 · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Improving vector space word representations using multilingual correlation
Manaal Faruqui and Chris Dyer. 2014 · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Shoou-I Yu, Lu Jiang, and Alexander Hauptmann. 2014 · 2014
Earlier work this paper cites.
Simple task-specific bilingual word embeddings
Stephan Gouws and Anders Søgaard. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nicholas Johnston, Andrew Rabinovich, and Kevin Murphy. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. 2015 · 2015
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien. 2016 · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Yuke Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, Stephanie Chen, Yannis Kalantidis, L. Li, D. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Cross-lingual image caption generation
Takashi Miyazaki and Nobuyuki Shimizu. 2016 · 2016
Cited alongside, same era.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Learning bilingual word embeddings with (almost) no bilingual data
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017 · 2017
Cited alongside, same era.
Cross-lingual character-level neural morphological tagging
Ryan Cotterell and Georg Heigold. 2017 · 2017
Cited alongside, same era.
Learning to parse and translate improves neural machine translation
Improving what cross-modal retrieval models learn through object-oriented inter- and intra-modal attention networks
Po-Yao Huang, Vaibhav, Xiaojun Chang, and Alexander G. Hauptmann. 2019d · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 2019
Later among the works it cites.
Use what you have: Video retrieval using representations from collaborative experts
Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Howto100M: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Akiko Eriguchi, Yoshimasa Tsuruoka, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Image pivoting for learning multilingual multimodal representations
Spandana Gella, Rico Sennrich, Frank Keller, and Mirella Lapata. 2017 · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Adversarial deep averaging networks for cross-lingual sentiment classification
Xilun Chen, Yu Sun, Ben Athiwaratkun, Claire Cardie, and Kilian Weinberger. 2018 · 2018
Cited alongside, same era.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Later among the works it cites.
Massively multilingual transfer for NER
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Later among the works it cites.
Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing
Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson. 2019 · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. 2019 · 2019
Later among the works it cites.
Fine-grained action retrieval through multiple parts-of-speech embeddings
Michael Wray, Diane Larlus, Gabriela Csurka, and Dima Damen. 2019 · 2019
Later among the works it cites.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
Shijie Wu and Mark Dredze. 2019b · 2019
Later among the works it cites.
Billion-scale semi-supervised learning for image classification
I. Zeki Yalniz, Hervé Jégou, Kan Chen, Manohar Paluri, and Dhruv Mahajan. 2019 · 2019
Later among the works it cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Later among the works it cites.
Learning to scale multilingual representations for vision-language tasks
Andrea Burns, Donghyun Kim, Derry Wijaya, Kate Saenko, and Bryan A. Plummer. 2020 · 2020
Later among the works it cites.
Forward and backward multimodal nmt for improved monolingual and multilingual cross-modal retrieval
Po-Yao Huang, Xiaojun Chang, Alexander Hauptmann, and Eduard Hovy. 2020a · 2020
Later among the works it cites.
Multi-modal dense video captioning
Vladimir Iashin and Esa Rahtu. 2020 · 2020
Later among the works it cites.
MULE: Multimodal Universal Language Embedding
Donghyun Kim, Kuniaki Saito, Kate Saenko, Stan Sclaroff, and Bryan A. Plummer. 2020 · 2020
Later among the works it cites.
TVQA+: Spatio-temporal grounding for video question answering
Jie Lei, Licheng Yu, Tamara Berg, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
MLQA: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020 · 2020
Later among the works it cites.
End-to-End Learning of Visual Representations from Uncurated Instructional Videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020 · 2020
Later among the works it cites.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Later among the works it cites.
Visual grounding in video for unsupervised word translation
Gunnar A. Sigurdsson, Jean-Baptiste Alayrac, Aida Nematzadeh, Lucas Smaira, Mateusz Malinowski, João Carreira, Phil Blunsom, and Andrew Zisserman. 2020 · 2020
Later among the works it cites.
Neural machine translation with universal visual representation
Zhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama, Eiichiro Sumita, Zuchao Li, and Hai Zhao. 2020 · 2020
Later among the works it cites.
Support-set bottlenecks for video-text representation learning
Mandela Patrick, Po-Yao Huang, Yuki M. Asano, Florian Metze, Alexander Hauptmann, João Henriques, and Andrea Vedaldi. 2021 · 2021
Closest in time.