Fetching the paper…
Reading the bibliography…
In this paper, we propose a new approach to learn multimodal multilingual embeddings for matching images and their relevant captions in two languages.
Sheep, goats, lambs and wolves a statistical analysis of speaker performance in the nist 1998 speaker recognition evaluation
George Doddington, Walter Liggett, Alvin Martin, Mark Przybocki, and Douglas Reynolds. 1998 · 1998
Earlier work this paper cites.
Applying conditional random fields to japanese morphological analysis
Taku Kudo, Kaoru Yamamoto, and Yuji Matsumoto. 2004 · 2004
Earlier work this paper cites.
Europarl: A Parallel Corpus for Statistical Machine Translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc Aurelio Ranzato, and Tomas Mikolov. 2013 · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Exploiting similarities among languages for machine translation
Tomas Mikolov, Quoc V. Le, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Chris D. Manning, and Andrew Y. Ng. 2013 · 2013
Earlier work this paper cites.
Improving zero-shot learning by mitigating the hubness problem
Georgiana Dinu and Marco Baroni. 2014 · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2014 · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li. 2014 · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L. Yuille. 2014 · 2014
Cited alongside, same era.
From image descriptions to visual denotations
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Cited alongside, same era.
Image-mediated learning for zero-shot cross-lingual document retrieval
Ruka Funaki and Hideki Nakayama. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni. 2015 · 2015
Cited alongside, same era.
Cross-lingual image caption generation
Takashi Miyazaki and Nobuyuki Shimizu. 2016 · 2016
Later among the works it cites.
Self-critical sequence training for image captioning
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel. 2016 · 2016
Later among the works it cites.
Bottom-up and top-down attention for image captioning and VQA
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2017 · 2017
Later among the works it cites.
Multilingual multi-modal embeddings for natural language processing
Iacer Calixto, Qun Liu, and Nick Campbell. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bridge correlational neural networks for multilingual multimodal representation learning
Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, and Balaraman Ravindran. 2015 · 2015
Cited alongside, same era.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. 2015 · 2015
Cited alongside, same era.
Automatic description generation from images: A survey of models, datasets, and evaluation measures
Raffaella Bernardi, Ruket Çakici, Desmond Elliott, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis, Frank Keller, Adrian Muscat, and Barbara Plank. 2016 · 2016
Cited alongside, same era.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Cited alongside, same era.
Multimodal pivots for image caption translation
Julian Hitschler and Stefan Riezler. 2016 · 2016
Cited alongside, same era.
Instance-aware image and sentence matching with selective multimodal LSTM
Yan Huang, Wei Wang, and Liang Wang. 2016 · 2016
Cited alongside, same era.
Knowing when to look: Adaptive attention via A visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2016 · 2016
Cited alongside, same era.
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2017 · 2017
Later among the works it cites.
VSE++: improved visual-semantic embeddings
Fartash Faghri, David J. Fleet, Ryan Kiros, and Sanja Fidler. 2017 · 2017
Later among the works it cites.
Image pivoting for learning multilingual multimodal representations
Spandana Gella, Rico Sennrich, Frank Keller, and Mirella Lapata. 2017 · 2017
Later among the works it cites.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2017 · 2017
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, and Svetlana Lazebnik. 2017 · 2017
Later among the works it cites.
STAIR captions: Constructing a large-scale japanese image caption dataset
Yuya Yoshikawa, Yutaro Shigeto, and Akikazu Takeuchi. 2017 · 2017
Later among the works it cites.
Loss in translation: Learning bilingual word mapping with a retrieval criterion
Armand Joulin, Piotr Bojanowski, Tomas Mikolov, Hervé Jégou, and Edouard Grave. 2018 · 2018
Later among the works it cites.
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018 · 2018
Later among the works it cites.