Fetching the paper…
Reading the bibliography…
Multimodal machine learning algorithms aim to learn visual-textual correspondences.
Recognition memory for nouns as a function of abstractness and frequency
Aloysia M. Gorman. 1961 · 1961
Earlier work this paper cites.
Differential context effects in the comprehension of abstract and concrete verbal materials
Paula J. Schwanenflugel and Edward J. Shoben. 1983 · 1983
Earlier work this paper cites.
Dual coding theory: Retrospect and current status
Allan Paivio. 1991 · 1991
Earlier work this paper cites.
Concrete words are easier to recall than abstract words: Evidence for a semantic contribution to short-term serial recall
Ian Walker and Charles Hulme. 1999 · 1999
Earlier work this paper cites.
What is hard to learn is easy to forget: The roles of word concreteness, cognate status, and word frequency in foreign-language vocabulary learning and forgetting
Annette De Groot and Rineke Keijzer. 2000 · 2000
Earlier work this paper cites.
Mallet: A machine learning for language toolkit
Andrew Kachites McCallum. 2002 · 2002
Earlier work this paper cites.
Matching words and pictures
Kobus Barnard, Pinar Duygulu, David Forsyth, Nando De Freitas, David M. Blei, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
Latent Dirichlet allocation
David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
Image annotation using SVM
Claudio Cusano, Gianluigi Ciocca, and Raimondo Schettini. 2004 · 2004
Earlier work this paper cites.
The University of South Florida free association, rhyme, and word fragment norms
Douglas L. Nelson, Cathy L. McEvoy, and Thomas A. Schreiber. 2004 · 2004
Earlier work this paper cites.
A discriminative kernel-based approach to rank images from text queries
David Grangier and Samy Bengio. 2008 · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
TVGraz: Multi-modal learning of object categories by combining textual and visual features
Inayatullah Khan, Amir Saffari, and Horst Bischof. 2009 · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. 2010 · 2010
Earlier work this paper cites.
New trends and ideas in visual concept detection: the MIR flickr retrieval evaluation initiative
Mark J. Huiskes, Bart Thomee, and Michael S. Lew. 2010 · 2010
Earlier work this paper cites.
Overview of the Wikipedia retrieval task at ImageCLEF 2010
Adrian Popescu, Theodora Tsikrika, and Jana Kludas. 2010 · 2010
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval
Nikhil Rasiwasia, Jose Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert R.G. Lanckriet, Roger Levy, and Nuno Vasconcelos. 2010 · 2010
Earlier work this paper cites.
Software framework for topic modelling with large corpora
Radim Řehůřek and Petr Sojka. 2010 · 2010
Earlier work this paper cites.
Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
Richard Socher and Li Fei-Fei. 2010 · 2010
Earlier work this paper cites.
Learning cross-modality similarity for multinomial data
Yangqing Jia, Mathieu Salzmann, and Trevor Darrell. 2011 · 2011
Cited alongside, same era.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordóñez, Girish Kulkarni, and Tamara L. Berg. 2011 · 2011
Cited alongside, same era.
Literal and metaphorical sense identification through concrete and abstract context
Peter D. Turney, Yair Neuman, Dan Assaf, and Yohai Cohen. 2011 · 2011
Cited alongside, same era.
Grounded models of semantic representation
Carina Silberer and Mirella Lapata. 2012 · 2012
Cited alongside, same era.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff A. Bilmes, and Karen Livescu. 2013 · 2013
Cited alongside, same era.
Fast image tagging
Minmin Chen, Alice X. Zheng, and Kilian Q. Weinberger. 2013 · 2013
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Later among the works it cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K. Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. Lawrence Zitnick, and Geoffrey Zweig. 2015 · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Later among the works it cites.
Image specificity
Mainak Jas and Devi Parikh. 2015 · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Concreteness and corpora: A theoretical and practical analysis
Felix Hill, Douwe Kiela, and Anna Korhonen. 2013 · 2013
Cited alongside, same era.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Cited alongside, same era.
Babytalk: Understanding and generating simple image descriptions
Girish Kulkarni, Visruth Premraj, Vicente Ordóñez, Sagnik Dhar, Siming Li, Yejin Choi, Alexander C. Berg, and Tamara L. Berg. 2013 · 2013
Cited alongside, same era.
A biterm topic model for short texts
Xiaohui Yan, Jiafeng Guo, Yanyan Lan, and Xueqi Cheng. 2013 · 2013
Cited alongside, same era.
Supervised coupled dictionary learning with group structures for multi-modal retrieval
Yue Ting Zhuang, Yan Fei Wang, Fei Wu, Yin Zhang, and Wei Ming Lu. 2013 · 2013
Cited alongside, same era.
On the role of correlation and abstraction in cross-modal multimedia retrieval
Jose Costa Pereira, Emanuele Coviello, Gabriel Doyle, Nikhil Rasiwasia, Gert R.G. Lanckriet, Roger Levy, and Nuno Vasconcelos. 2014 · 2014
Cited alongside, same era.
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel. 2015 · 2015
Later among the works it cites.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni. 2015 · 2015
Later among the works it cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Later among the works it cites.
Digitised books
British Library Labs. 2016 · 2016
Later among the works it cites.
Under the hood: Building accessibility tools for the visually impaired on facebook
Dario Garcia Garcia, Manohar Paluri, and Shaomei Wu. 2016 · 2016
Later among the works it cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Later among the works it cites.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache. 2016 · 2016
Later among the works it cites.
Visually grounded meaning representations
Carina Silberer, Vittorio Ferrari, and Mirella Lapata. 2016 · 2016
Later among the works it cites.
Cross-modal retrieval with cnn visual features: A new baseline
Yunchao Wei, Yao Zhao, Canyi Lu, Shikui Wei, Luoqi Liu, Zhenfeng Zhu, and Shuicheng Yan. 2016 · 2016
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2017 · 2017
Later among the works it cites.
Exploring multi-modal text+image models to distinguish between abstract and concrete nouns
Sai Abishek Bhaskar, Maximilian Köper, Sabine Schulte im Walde, and Diego Frassinelli. 2017 · 2017
Later among the works it cites.
OpenImages: A public dataset for large-scale multi-label and multi-class image classification
Ivan Krasin, Tom Duerig, Neil Alldrin, Vittorio Ferrari, Sami Abu-El-Haija, Alina Kuznetsova, Hassan Rom, Jasper Uijlings, Stefan Popov, Andreas Veit, Serge Belongie, Victor Gomes, Abhinav Gupta, Chen Sun, Gal Chechik, David Cai, Zheyun Feng, Dhyanesh Narayanan, and Kevin Murphy. 2017 · 2017
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher. 2017 · 2017
Later among the works it cites.
Learning from noisy large-scale datasets with minimal supervision
Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin, Abhinav Gupta, and Serge Belongie. 2017 · 2017
Later among the works it cites.