Fetching the paper…
Reading the bibliography…
Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors.
H. Hotelling, “Relations Between Two Sets of Variates,” Biometrika , 1936
1936
Earlier work this paper cites.
H. McGurk and J. Macdonald, “Hearing lips and seeing voices.” Nature , 1976
1976
Earlier work this paper cites.
J. B. Kruskal, “An Overview of Sequence Comparison: Time Warps, String Edits, and Macromolecules,” Society for Industrial and Applied Mathematics Review , vol. 25, no. 2, pp. 201–237, 1983
1983
Earlier work this paper cites.
B. P. Yuhas, M. H. Goldstein, and T. J. Sejnowski, “Integration of Acoustic and Visual Speech Signals Using Neural Networks,” IEEE Communications Magazine , 1989
1989
Earlier work this paper cites.
B. H. Juang and L. R. Rabiner, “Hidden Markov Models for Speech Recognition,” Technometrics , 1991
1991
Earlier work this paper cites.
P. F. Brown, S. A. D. Pietra, V. J. D. Pietra, and R. L. Mercer, “The mathematics of statistical machine translation: Parameter estimation,” Computational linguistics , pp. 263–311, 1993
1993
Earlier work this paper cites.
G. Hinton and R. S. Zemel, “Autoencoders, minimum description length and Helmoltz free energy,” in NIPS , 1993
1993
Earlier work this paper cites.
P. Cosi, E. Caldognetto, K. Vagges, G. Mian, M. Contolini, C. per Le Ricerche, and C. di Fonetica, “Bimodal recognition experiments with recurrent neural networks,” in ICASSP , 1994
1994
Earlier work this paper cites.
H. Bourlard and S. Dupont, “A mew ASR approach based on independent processing and recombination of partial frequency bands,” in International Conference on Spoken Language , 1996
1996
Earlier work this paper cites.
A. Hunt and A. W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” ICASSP , 1996
1996
Earlier work this paper cites.
S. Vogel, H. Ney, and C. Tillmann, “HMM-based word alignment in statistical translation,” in Computational Linguistics , 1996
1996
Earlier work this paper cites.
M. Brand, N. Oliver, and A. Pentland, “Coupled hidden Markov models for complex action recognition,” CVPR , 1997
1997
Earlier work this paper cites.
C. Bregler, M. Covell, and M. Slaney, “Video rewrite: Driving visual speech with audio,” in SIGGRAPH , 1997
1997
Earlier work this paper cites.
Z. Ghahramani and M. I. Jordan, “Factorial hidden Markov models,” Machine Learning , 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , 1997
1997
Earlier work this paper cites.
A. Blum and T. Mitchell, “Combining labeled and unlabeled data with co-training,” Computational learning theory , 1998
1998
Earlier work this paper cites.
R. W. Lienhart, “Comparison of automatic shot boundary detection algorithms,” Proceedings of SPIE , 1998
1998
Earlier work this paper cites.
T. Masuko, T. Kobayashi, M. Tamura, J. Masubuchi, and K. Tokuda, “Text-to-Visual Speech Synthesis Based on Parameter Generation from HMM,” in ICASSP , 1998
1998
Earlier work this paper cites.
P. L. Lai and C. Fyfe, “Kernel and nonlinear canonical correlation analysis,” International Journal of Neural Systems , 2000
2000
Earlier work this paper cites.
A. Ratnaparkhi, “Trainable methods for surface natural language generation,” in NAACL , 2000
2000
Earlier work this paper cites.
M. Slaney and M. Covell, “FaceSync: A linear operator for measuring synchronization of video facial images and audio tracks,” in NIPS , 2000
2000
Earlier work this paper cites.
B. Coyne and R. Sproat, “WordsEye: an automatic text-to-scene conversion system,” in SIGGRAPH , 2001
2001
Earlier work this paper cites.
J. D. Lafferty, A. McCallum, and F. C. N. Pereira, “Conditional Random Fields : Probabilistic Models for Segmenting and Labeling Sequence Data,” in ICML , 2001
2001
Earlier work this paper cites.
A. Sarkar, “Applying Co-Training methods to statistical parsing,” in ACL , 2001
2001
Earlier work this paper cites.
A. Kojima, T. Tamura, and K. Fukunaga, “Natural language description of human activities from video images based on concept hierarchy of actions,” IJCV , 2002
2002
Earlier work this paper cites.
A. V. Nefian, L. Liang, X. Pi, L. Xiaoxiang, C. Mao, and K. Murphy, “A coupled HMM for audio-visual speech recognition,” Interspeech , vol. 2, 2002
2002
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-j. Zhu, “BLEU: a Method for Automatic Evaluation of Machine Translation,” ACL , 2002
2002
Earlier work this paper cites.
K. Barnard, P. Duygulu, D. Forsyth, N. de Freitas, D. M. Blei, and M. I. Jordan, “Matching Words and Pictures,” JMLR , 2003
2003
Earlier work this paper cites.
A. Garg, V. Pavlovic, and J. M. Rehg, “Boosted learning in dynamic bayesian networks for multimodal speaker detection,” Proceedings of the IEEE , 2003
2003
Earlier work this paper cites.
D. R. Hardoon, S. Szedmak, and J. Shawe-taylor, “Canonical correlation analysis; An overview with application to learning methods,” Tech. Rep., 2003
2003
Earlier work this paper cites.
A. Levin, P. Viola, and Y. Freund, “Unsupervised improvement of visual detectors using cotraining,” in ICCV , 2003
2003
Earlier work this paper cites.
C.-Y. Lin and E. Hovy, “Automatic Evaluation of Summaries Using N-gram Co-Occurrence Statistics,” NAACL , 2003
2003
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audio-visual speech,” Proceedings of the IEEE , 2003
2003
Earlier work this paper cites.
K. Sjölander, “An HMM-based system for automatic segmentation and alignment of speech,” in Proceedings of Fonetik , 2003
2003
Earlier work this paper cites.
M. A. Krogel and T. Scheffer, “Multi-relational learning, text mining, and semi-supervised learning for functional genomics,” Machine Learning , 2004
2004
Earlier work this paper cites.
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” IJCV , 2004
2004
Earlier work this paper cites.
C. Yu and D. Ballard, “On the Integration of Grounding Language and Lear ning Objects,” in AAAI , 2004
2004
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, and P. Wellner, “The AMI Meeting Corpus: A Pre-Announcement,” in Int. Conf. on Methods and Techniques in Behavioral Research , 2005
2005
Earlier work this paper cites.
C. G. M. Snoek and M. Worring, “Multimodal video indexing: A review of the state-of-the-art,” Multimedia Tools and Applications , 2005
2005
Earlier work this paper cites.
Z. Wu, L. Cai, and H. Meng, “Multi-level Fusion of Audio and Visual Features for Speaker Identification,” Advances in Biometrics , 2005
2005
Earlier work this paper cites.
C. M. Christoudias, K. Saenko, L.-P. Morency, and T. Darrell, “Co-Adaptation of audio-visual speech and gesture classifiers,” in ICMI , 2006
2006
Earlier work this paper cites.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A Fast Learning Algorithm for Deep Belief Nets,” Neural Computation , 2006
2006
Earlier work this paper cites.
C. Sutton and A. McCallum, “Introduction to Conditional Random Fields for Relational Learning,” in Introduction to Statistical Relational Learning . MIT Press, 2006
2006
Earlier work this paper cites.
Dynamic Time Warping . Berlin, Heidelberg: Springer Berlin Heidelberg, 2007, pp. 69–84
2007
Earlier work this paper cites.
A. Haubold and J. R. Kender, “Alignment of speech to highly imperfect text transcriptions,” in ICME , 2007
2007
Earlier work this paper cites.
A. Quattoni, S. Wang, L.-P. Morency, M. Collins, and T. Darrell, “Hidden conditional random fields.” IEEE TPAMI , vol. 29, 2007
2007
Earlier work this paper cites.
S. Reiter, B. Schuller, and G. Rigoll, “Hidden Conditional Random Fields for Meeting Segmentation,” ICME , 2007
2007
Earlier work this paper cites.
M. E. Sargin, Y. Yemez, E. Erzin, and A. M. Tekalp, “Audiovisual synchronization and fusion using canonical correlation analysis,” IEEE Trans. Multimedia , 2007
2007
Earlier work this paper cites.
L. W. Barsalou, “Grounded cognition,” Annual review of psychology , 2008
2008
Earlier work this paper cites.
G. Castellano, L. Kessous, and G. Caridakis, “Emotion recognition through multiple modalities: Face, body gesture, speech,” LNCS , 2008
2008
Earlier work this paper cites.
C. M. Christoudias, R. Urtasun, and T. Darrell, “Multi-view learning in the presence of view disagreement,” in UAI , 2008
2008
Earlier work this paper cites.
T. Cour, C. Jordan, E. Miltsakaki, and B. Taskar, “Movie / Script : Alignment and Parsing of Video and Text Transcription,” in ECCV , 2008, pp. 1–14
2008
Earlier work this paper cites.
M. Gurban, J.-P. Thiran, T. Drugman, and T. Dutoit, “Dynamic Modality Weighting for Multi-stream HMMs in Audio-Visual Speech Recognition,” in ICMI , 2008
2008
Earlier work this paper cites.
T. Qin, T.-y. Liu, X.-d. Zhang, D.-s. Wang, and H. Li, “Global Ranking Using Continuous Conditional Random Fields,” in NIPS , 2008
2008
Earlier work this paper cites.
S. Deena and A. Galata, “Speech-Driven Facial Animation Using a Shared Gaussian Process Latent Variable Model,” in Advances in Visual Computing , 2009
2009
Earlier work this paper cites.
A. Farhadi, I. Endres, D. Hoiem, and D. Forsyth, “Describing objects by their attributes,” in CVPR , 2009
2009
Earlier work this paper cites.
P. Gehler and S. Nowozin, “On Feature Combination for Multiclass Object Classification,” in ICCV , 2009
2009
Earlier work this paper cites.
M. Palatucci, G. E. Hinton, D. Pomerleau, and T. M. Mitchell, “Zero-Shot Learning with Semantic Output Codes,” in NIPS , 2009
2009
Earlier work this paper cites.
R. Salakhutdinov and G. E. Hinton, “Deep Boltzmann Machines,” in International conference on artificial intelligence and statistics , 2009
2009
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” Speech Communication , vol. 51, 2009
2009
Earlier work this paper cites.
Z. Zeng, M. Pantic, G. I. Roisman, and T. S. Huang, “A Survey of Affect Recognition Methods: Audio, Visual, and Spontaneous Expressions,” IEEE TPAMI , 2009
2009
Earlier work this paper cites.
F. Zhou and F. Torre, “Canonical time warping for alignment of human behavior,” in NIPS , 2009
2009
Earlier work this paper cites.
P. K. Atrey, M. A. Hossain, A. El Saddik, and M. S. Kankanhalli, “Multimodal fusion for multimedia analysis: A survey,” 2010
2010
Earlier work this paper cites.
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh, “VizWiz: Nearly Real-Time Answers to Vvisual Questions,” in UIST , 2010
2010
Earlier work this paper cites.
M. M. Bronstein, A. M. Bronstein, F. Michel, and N. Paragios, “Data Fusion through Cross-modality Metric Learning using Similarity-Sensitive Hashing,” in CVPR , 2010
2010
Earlier work this paper cites.
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth, “Every picture tells a story: Generating sentences from images,” LNCS , 2010
2010
Earlier work this paper cites.
Y. Feng and M. Lapata, “Visual Information in Semantic Representation,” in NAACL , 2010
2010
Earlier work this paper cites.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in International Conference on Artificial Intelligence and Statistics , 2010
2010
Earlier work this paper cites.
M. M. Khapra, A. Kumaran, and P. Bhattacharyya, “Everybody loves a rich cousin: An empirical study of transliteration through bridge languages,” in NAACL , 2010
2010
Earlier work this paper cites.
G. McKeown, M. F. Valstar, R. Cowie, and M. Pantic, “The SEMAINE corpus of emotionally coloured character interactions,” in IEEE International Conference on Multimedia and Expo , 2010
2010
Earlier work this paper cites.
N. Rasiwasia, J. Costa Pereira, E. Coviello, G. Doyle, G. R. Lanckriet, R. Levy, and N. Vasconcelos, “A new approach to cross-modal multimedia retrieval,” in ACMMM , 2010
2010
Earlier work this paper cites.
R. Socher and L. Fei-Fei, “Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,” in CVPR , 2010
2010
Earlier work this paper cites.
J. Weston, S. Bengio, and N. Usunier, “Web Scale Image Annotation: Learning to Rank with Joint Word-Image Embeddings Image Annotation,” ECML , 2010
2010
Earlier work this paper cites.
M. Wöllmer, A. Metallinou, F. Eyben, B. Schuller, and S. Narayanan, “Context-Sensitive Multimodal Emotion Recognition from Speech and Facial Expression using Bidirectional LSTM Modeling,” INTERSPEECH , 2010
2010
Earlier work this paper cites.
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S. C. Zhu, “I2T: Image parsing to text description,” Proceedings of the IEEE , 2010
2010
Earlier work this paper cites.
F. De la Torre and J. F. Cohn, “Facial Expression Analysis,” in Guide to Visual Analysis of Humans: Looking at People , 2011
2011
Earlier work this paper cites.
M. Glodek, S. Tschechne, G. Layher, M. Schels, T. Brosch, S. Scherer, M. Kächele, M. Schmidt, H. Neumann, G. Palm, and F. Schwenker, “Multiple classifier systems for the classification of audio-visual emotional states,” LNCS , 2011
2011
Earlier work this paper cites.
M. Gönen and E. Alpaydın, “Multiple Kernel Learning Algorithms,” JMLR , 2011
2011
Earlier work this paper cites.
S. Kumar and R. Udupa, “Learning hash functions for cross-view similarity search,” in IJCAI , 2011
2011
Earlier work this paper cites.
S. Li, G. Kulkarni, T. Berg, A. Berg, and Y. Choi, “Composing simple image descriptions using web-scale n-grams,” in CoNLL , 2011
2011
Earlier work this paper cites.
M. M. Louwerse, “Symbol interdependency in symbolic and embodied cognition,” Topics in Cognitive Science , 2011
2011
Earlier work this paper cites.
B. McFee and G. R. G. Lanckriet, “Learning Multi-modal Similarity,” JMLR , 2011
2011
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal Deep Learning,” ICML , 2011
2011
Earlier work this paper cites.
M. A. Nicolaou, H. Gunes, and M. Pantic, “Continuous Prediction of Spontaneous Affect from Multiple Cues and Modalities in Valence – Arousal Space,” IEEE TAC , 2011
2011
Earlier work this paper cites.
V. Ordonez, G. Kulkarni, and T. L. Berg, “Im2text: Describing images using 1 million captioned photographs,” in NIPS , 2011
2011
Cited alongside, same era.
G. A. Ramirez, T. Baltrušaitis, and L.-P. Morency, “Modeling Latent Discriminative Dynamic of Multi-Dimensional Affective Signals,” in ACII workshops , 2011
2011
Cited alongside, same era.
B. Schuller, M. F. Valstar, F. Eyben, G. McKeown, R. Cowie, and M. Pantic, “AVEC 2011 – The First International Audio / Visual Emotion Challenge,” in ACII , 2011
2011
Cited alongside, same era.
S. Shariat and V. Pavlovic, “Isotonic CCA for sequence alignment and activity recognition,” in ICCV , 2011
2011
Cited alongside, same era.
——, “WSABIE: Scaling up to large vocabulary image annotation,” in IJCAI , 2011
2011
Cited alongside, same era.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh, “VQA: Visual question answering,” in ICCV , 2015
2015
Later among the works it cites.
P. Bojanowski, R. Lajugie, E. Grave, F. Bach, I. Laptev, J. Ponce, and C. Schmid, “Weakly-Supervised Alignment of Video With Text,” in ICCV , 2015
2015
Later among the works it cites.
S. Chen and Q. Jin, “Multi-modal Dimensional Emotion Recognition Using Recurrent Neural Networks,” in Proceedings of the 5th International Workshop on Audio/Visual Emotion Challenge , 2015
2015
Later among the works it cites.
X. Chen, H. Fang, T.-Y. Lin, R. Vedantam, S. Gupta, P. Dollar, and L. Zitnick, “Microsoft COCO Captions: Data Collection and Evaluation Server,” 2015
2015
Later among the works it cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in NIPS , 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Yang, C. L. Teo, H. Daume, and Y. Aloimonos, “Corpus-Guided Sentence Generation of Natural Images,” in EMNLP , 2011
2011
Cited alongside, same era.
C. N. Anagnostopoulos, T. Iliou, and I. Giannoukos, “Features and classifiers for emotion recognition from speech: a survey from 2000 to 2011,” Artificial Intelligence Review , 2012
2012
Cited alongside, same era.
A. Barbu, A. Bridge, Z. Burchill, D. Coroian, S. Dickinson, S. Fidler, A. Michaux, S. Mussman, S. Narayanaswamy, D. Salvi, L. Schmidt, J. Shangguan, J. M. Siskind, J. Waggoner, S. Wang, J. Wei, Y. Yin, and Z. Zhang, “Video In Sentences Out,” in Proc. of the Conference on Uncertainty in Artificial Intelligence , 2012
2012
Cited alongside, same era.
E. Bruni, G. Boleda, M. Baroni, and N.-K. Tran, “Distributional Semantics in Technicolor,” in ACL , 2012
2012
Cited alongside, same era.
A. Gupta, Y. Verma, and C. V. Jawahar, “Choosing Linguistics over Vision to Describe Images,” in AAAI , 2012
2012
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury, “Deep Neural Networks for Acoustic Modeling in Speech Recognition,” IEEE Signal Processing Magazine , 2012
2012
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” NIPS , 2012
2012
Cited alongside, same era.
2015
Later among the works it cites.
J. Devlin, H. Cheng, H. Fang, S. Gupta, L. Deng, X. He, G. Zweig, and M. Mitchell, “Language Models for Image Captioning: The Quirks and What Works,” ACL , 2015
2015
Later among the works it cites.
S. K. D’mello and J. Kory, “A Review and Meta-Analysis of Multimodal Affect Detection Systems,” ACM Computing Surveys , 2015
2015
Later among the works it cites.
F. Feng, R. Li, and X. Wang, “Deep correspondence restricted Boltzmann machine for cross-modal retrieval,” Neurocomputing , 2015
2015
Later among the works it cites.
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu, “Are you talking to a machine? dataset and methods for multilingual image question answering,” NIPS , 2015
2015
Later among the works it cites.
N. Jaques, S. Taylor, A. Sano, and R. Picard, “Multi-task , Multi-Kernel Learning for Estimating Individual Wellbeing,” in Multimodal Machine Learning Workshop in conjunction with NIPS , 2015
2015
Later among the works it cites.
X. Jia, E. Gavves, B. Fernando, and T. Tuytelaars, “Guiding the Long-Short Term Memory Model for Image Caption Generation,” ICCV , 2015
2015
Later among the works it cites.
X. Jiang, F. Wu, Y. Zhang, S. Tang, W. Lu, and Y. Zhuang, “The classification of multi-modal data with hidden conditional random field,” Pattern Recognition Letters , 2015
2015
Later among the works it cites.
S. E. Kahou, X. Bouthillier, P. Lamblin, C. Gulchere, V. Michalski, K. Konda, J. Sebastien, P. Froumenty, Y. Dauphin, N. Boulanger-Lewandowski, R. C. Ferrari, M. Mirza, D. Warde-Farley, A. Courville, P. Vincent, R. Memisevic, C. Pal, and Y. Bengio, “EmoNets: Multimodal deep learning approaches for emotion recognition in video,” Journal on Multimodal User Interfaces , 2015
2015
Later among the works it cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in CVPR , 2015
2015
Later among the works it cites.
D. Kiela, L. Bulat, and S. Clark, “Grounding Semantics in Olfactory Perception,” in ACL , 2015
2015
Later among the works it cites.
D. Kiela and S. Clark, “Multi- and Cross-Modal Semantics Beyond Vision: Grounding in Auditory Perception,” EMNLP , 2015
2015
Later among the works it cites.
B. Klein, G. Lev, G. Sadeh, and L. Wolf, “Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation,” in CVPR , 2015
2015
Later among the works it cites.
R. Lebret, P. O. Pinheiro, and R. Collobert, “Phrase-based Image Captioning,” ICML , 2015
2015
Later among the works it cites.
Y. Li, S. Wang, Q. Tian, and X. Ding, “A survey of recent advances in visual feature detection,” Neurocomputing , 2015
2015
Later among the works it cites.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in ICCV , 2015
2015
Later among the works it cites.
J. Malmaud, J. Huang, V. Rathod, N. Johnston, A. Rabinovich, and K. Murphy, “What’s cookin’? interpreting cooking videos using text, speech and vision,” NAACL , 2015
2015
Later among the works it cites.
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille, “Deep Captioning with multimodal recurrent neural networks (m-RNN),” ICLR , 2015
2015
Later among the works it cites.
S. Moon, S. Kim, and H. Wang, “Multimodal Transfer Deep Learning for Audio-Visual Recognition,” NIPS Workshops , 2015
2015
Later among the works it cites.
Y. Mroueh, E. Marcheret, and V. Goel, “Deep multimodal learning for Audio-Visual Speech Recognition,” in ICASSP , 2015
2015
Later among the works it cites.
I. Naim, Y. Song, Q. Liu, L. Huang, H. Kautz, J. Luo, and D. Gildea, “Discriminative unsupervised alignment of natural language instructions with corresponding video segments,” in NAACL , 2015
2015
Later among the works it cites.
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik, “Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models,” in ICCV , 2015
2015
Later among the works it cites.
S. Poria, E. Cambria, and A. Gelbukh, “Deep Convolutional Neural Network Textual Features and Multiple Kernel Learning for Utterance-level Multimodal Sentiment Analysis,” EMNLP , 2015
2015
Later among the works it cites.
J. Rajendran, M. M. Khapra, S. Chandar, and B. Ravindran, “Bridge Correlational Neural Networks for Multilingual Multimodal Representation Learning,” in NAACL , 2015
2015
Later among the works it cites.
A. Rohrbach, M. Rohrbach, and B. Schiele, “The long-short story of movie description,” in Pattern Recognition , 2015
2015
Later among the works it cites.
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in ICLR , 2015
2015
Later among the works it cites.
M. Tapaswi, M. Bäuml, and R. Stiefelhagen, “Aligning plot synopses to videos for story-based retrieval,” IJMIR , 2015
2015
Later among the works it cites.
——, “Book2Movie: Aligning video scenes with book chapters,” in CVPR , 2015
2015
Later among the works it cites.
A. Torabi, C. Pal, H. Larochelle, and A. Courville, “Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research,” 2015
2015
Later among the works it cites.
R. Vedantam, C. L. Zitnick, and D. Parikh, “CIDEr: Consensus-based Image Description Evaluation Ramakrishna Vedantam,” in CVPR , 2015
2015
Later among the works it cites.
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko, “Translating Videos to Natural Language Using Deep Recurrent Neural Networks,” NAACL , 2015
2015
Later among the works it cites.
——, “Show and tell: A neural image caption generator,” in CVPR , 2015
2015
Later among the works it cites.
D. Wang, P. Cui, M. Ou, and W. Zhu, “Deep Multimodal Hashing with Orthogonal Regularization,” in IJCAI , 2015
2015
Later among the works it cites.
W. Wang, R. Arora, K. Livescu, and J. Bilmes, “On deep multi-view representation learning,” in ICML , 2015
2015
Later among the works it cites.
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” ICML , 2015
2015
Later among the works it cites.
R. Xu, C. Xiong, W. Chen, and J. J. Corso, “Jointly modeling deep video and compositional text to bridge vision and language in a unified framework,” in AAAI , 2015
2015
Later among the works it cites.
S. Yagcioglu, E. Erdem, A. Erdem, and R. Cakici, “A Distributed Representation Based Query Expansion Approach for Image Captioning,” in ACL , 2015
2015
Later among the works it cites.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, “Describing videos by exploiting temporal structure,” in CVPR , 2015
2015
Later among the works it cites.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books,” in ICCV , 2015
2015
Later among the works it cites.
“YouTube statistics,” https://www.youtube.com/yt/press/statistics.html (accessed Sept. 2016), accessed: 2016-09-30
2016
Later among the works it cites.
A. Agrawal, D. Batra, and D. Parikh, “Analyzing the Behavior of Visual Question Answering Models,” in EMNLP , 2016
2016
Later among the works it cites.
M. Baroni, “Grounding Distributional Semantics in the Visual World Grounding Distributional Semantics in the Visual World,” Language and Linguistics Compass , 2016
2016
Later among the works it cites.
R. Bernardi, R. Cakici, D. Elliott, A. Erdem, E. Erdem, N. Ikizler-Cinbis, F. Keller, A. Muscat, and B. Plank, “Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures,” JAIR , 2016
2016
Later among the works it cites.
Y. Cao, M. Long, J. Wang, Q. Yang, and P. S. Yu, “Deep Visual-Semantic Hashing for Cross-Modal Retrieval,” in KDD , 2016
2016
Later among the works it cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, Attend, and Spell: a Neural Network for Large Vocabulary Conversational Speech Recognition,” in ICASSP , 2016
2016
Later among the works it cites.
R. Collobert, C. Puhrsch, and G. Synnaeve, “Wav2Letter: an End-to-End ConvNet-based Speech Recognition System,” 2016
2016
Later among the works it cites.
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach, “Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding,” in EMNLP , 2016
2016
Later among the works it cites.
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, K. Saenko, and T. Darrell, in CVPR , 2016
2016
Later among the works it cites.
R. Hu, H. Xu, M. Rohrbach, J. Feng, K. Saenko, and T. Darrell, “Natural Language Object Retrieval,” in CVPR , 2016
2016
Later among the works it cites.
T.-H. K. Huang, F. Ferraro, N. Mostafazadeh, I. Misra, A. Agrawal, J. Devlin, R. Girshick, X. He, P. Kohli, D. Batra et al. , “Visual storytelling.” NAACL, 2016
2016
Later among the works it cites.
Q. Jin and J. Liang, “Video Description Generation using Audio and Visual Cues,” in ICMR , 2016
2016
Later among the works it cites.
J. Lu, J. Yang, D. Batra, and D. Parikh, “Hierarchical Co-Attention for Visual Question Answering,” in NIPS , 2016
2016
Later among the works it cites.
B. Mahasseni and S. Todorovic, “Regularizing Long Short Term Memory with 3D Human-Skeleton Sequences for Action Recognition,” in CVPR , 2016
2016
Later among the works it cites.
E. Mansimov, E. Parisotto, J. L. Ba, and R. Salakhutdinov, “Generating Images from Captions with Attention,” in ICLR , 2016
2016
Later among the works it cites.
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy, “Generation and Comprehension of Unambiguous Object Descriptions,” in CVPR , 2016
2016
Later among the works it cites.
H. Mei, M. Bansal, and M. R. Walter, “Listen, attend, and walk: Neural mapping of navigational instructions to action sequences,” AAAI , 2016
2016
Later among the works it cites.
N. Neverova, C. Wolf, G. Taylor, and F. Nebout, “ModDrop: Adaptive multi-modal gesture recognition,” IEEE TPAMI , 2016
2016
Later among the works it cites.
B. Nojavanasghari, D. Gopinath, J. Koushik, T. Baltrušaitis, and L.-P. Morency, “Deep multimodal fusion for persuasiveness prediction,” in ICMI , 2016
2016
Later among the works it cites.
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman, “Visually Indicated Sounds,” in CVPR , 2016
2016
Later among the works it cites.
Y. Pan, T. Mei, T. Yao, H. Li, and Y. Rui, “Jointly Modeling Embedding and Translation to Bridge Video and Language,” in CVPR , 2016
2016
Later among the works it cites.
S. S. Rajagopalan, L.-P. Morency, T. Baltrušaitis, and R. Goecke, “Extending Long Short-Term Memory for Multi-View Structured Learning,” ECCV , 2016
2016
Later among the works it cites.
S. Reed, Z. Akata, X. Yan, L. Logeswaran, H. Lee, and B. Schiele, “Generative Adversarial Text to Image Synthesis,” in ICML , 2016
2016
Later among the works it cites.
E. Shutova, D. Kelia, and J. Maillard, “Black Holes and White Rabbits : Metaphor Identification with Visual Features,” NAACL , 2016
2016
Later among the works it cites.
Y. C. Song, I. Naim, A. A. Mamun, K. Kulkarni, P. Singla, J. Luo, D. Gildea, and H. Kautz, “Unsupervised Alignment of Actions in Video with Text Descriptions,” in IJCAI , 2016
2016
Later among the works it cites.
G. Trigeorgis, F. Ringeval, R. Brueckner, E. Marchi, M. A. Nicolaou, B. Schuller, and S. Zafeiriou, “Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network,” in ICASSP , 2016
2016
Later among the works it cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A Generative Model for Raw Audio,” 2016
2016
Later among the works it cites.
A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu, “Pixel Recurrent Neural Networks,” ICML , 2016
2016
Later among the works it cites.
I. Vendrov, R. Kiros, S. Fidler, and R. Urtasun, “Order-Embeddings of Images and Language,” in ICLR , 2016
2016
Later among the works it cites.
L. Wang, Y. Li, and S. Lazebnik, “Learning Deep Structure-Preserving Image-Text Embeddings,” in CVPR , 2016
2016
Later among the works it cites.
C. Xiong, S. Merity, and R. Socher, “Dynamic memory networks for visual and textual question answering,” ICML , 2016
2016
Later among the works it cites.
H. Xu and K. Saenko, “Ask, attend and answer: Exploring question-guided spatial attention for visual question answering,” ECCV , 2016
2016
Later among the works it cites.
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola, “Stacked Attention Networks for Image Question Answering,” in CVPR , 2016
2016
Later among the works it cites.
H. Yu, J. Wang, Z. Huang, Y. Yang, and W. Xu, “Video paragraph captioning using hierarchical recurrent neural networks,” CVPR , 2016
2016
Later among the works it cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling Context in Referring Expressions,” in ECCV , 2016
2016
Later among the works it cites.
H. Zhang, Z. Hu, Y. Deng, M. Sachan, Z. Yan, and E. P. Xing, “Learning Concept Taxonomies from Multi-modal Data,” in ACL , 2016
2016
Later among the works it cites.
“TRECVID Multimedia Event Detection 2011 Evaluation,” https://www.nist.gov/multimodal-information-group/trecvid-multimedia-event-detection-2011-evaluation , accessed: 2017-01-21
2017
Closest in time.
I. D. Gebru, S. Ba, X. Li, and R. Horaud, “Audio-visual speaker diarization based on spatiotemporal bayesian fusion,” TPAMI , 2017
2017
Closest in time.
Q.-y. Jiang and W.-j. Li, “Deep Cross-Modal Hashing,” in CVPR , 2017
2017
Closest in time.
A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele, “Movie description,” International Journal of Computer Vision , 2017
2017
Closest in time.
K.-H. Zeng, T.-H. Chen, C.-Y. Chuang, Y.-H. Liao, J. C. Niebles, and M. Sun, “Leveraging Video Descriptions to Learn Video Question Answering,” in AAAI , 2017
2017
Closest in time.