Fetching the paper…
Reading the bibliography…
Combining complementary information from multiple modalities is intuitively appealing for improving the performance of learning-based approaches.
Factorial hidden markov models
Z. Ghahramani and M. I. Jordan · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
A coupled hmm for audio-visual speech recognition
A. V. Nefian, L. Liang, X. Pi, L. Xiaoxiang, C. Mao, and K. Murphy · 2002
Earlier work this paper cites.
Affect recognition from face and body: early fusion vs. late fusion
H. Gunes and M. Piccardi · 2005
Earlier work this paper cites.
Early versus late fusion in semantic video analysis
C. G. Snoek, M. Worring, and A. W. Smeulders · 2005
Earlier work this paper cites.
Dynamic modality weighting for multi-stream hmms inaudio-visual speech recognition
M. Gurban, J.-P. Thiran, T. Drugman, and T. Dutoit · 2008
Earlier work this paper cites.
On feature combination for multiclass object classification
P. Gehler and S. Nowozin · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Multimodal fusion for multimedia analysis: a survey
P. K. Atrey, M. A. Hossain, A. El Saddik, and M. S. Kankanhalli · 2010
Earlier work this paper cites.
Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional lstm modeling
M. Wöllmer, A. Metallinou, F. Eyben, B. Schuller, and S. Narayanan · 2010
Earlier work this paper cites.
Multiple classifier systems for the classification of audio-visual emotional states
M. Glodek, S. Tschechne, G. Layher, M. Schels, T. Brosch, S. Scherer, M. Kächele, M. Schmidt, H. Neumann, G. Palm, et al · 2011
Earlier work this paper cites.
Multiple kernel learning algorithms
M. Gönen and E. Alpaydın · 2011
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Modeling latent discriminative dynamic of multi-dimensional affective signals
G. A. Ramirez, T. Baltrušaitis, and L.-P. Morency · 2011
Earlier work this paper cites.
Similarity component analysis
S. Changpinyo, K. Liu, and F. Sha · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
A. Graves, A.-r. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
M. Lin, Q. Chen, and S. Yan · 2013
Cited alongside, same era.
Multiple object recognition with visual attention
J. Ba, V. Mnih, and K. Kavukcuoglu · 2014
Cited alongside, same era.
Searching for exotic particles in high-energy physics with deep learning
P. Baldi, P. Sadowski, and D. Whiteson · 2014
Cited alongside, same era.
Multiple kernel learning for visual object recognition: A review
S. S. Bucak, R. Jin, and A. K. Jain · 2014
Cited alongside, same era.
Multi-scale recognition with dag-cnns
S. Yang and D. Ramanan · 2015
Later among the works it cites.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun · 2015
Later among the works it cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
W. Chan, N. Jaitly, Q. Le, and O. Vinyals · 2016
Later among the works it cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Parallel recurrent neural network architectures for feature-rich session-based recommendations
B. Hidasi, M. Quadrana, A. Karatzoglou, and D. Tikk · 2016
Later among the works it cites.
Video description generation using audio and visual cues
Q. Jin and J. Liang · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Kim · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, et al · 2014
Cited alongside, same era.
Majority vote of diverse classifiers for late fusion
E. Morvant, A. Habrard, and S. Ayache · 2014
Cited alongside, same era.
Multi-source deep learning for human pose estimation
W. Ouyang, X. Chu, and X. Wang · 2014
Cited alongside, same era.
Are you talking to a machine? dataset and methods for multilingual image question
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Cited alongside, same era.
Hypercolumns for object segmentation and fine-grained localization
B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik · 2015
Cited alongside, same era.
Collaborative layer-wise discriminative learning in deep neural networks
X. Jin, Y. Chen, J. Dong, J. Feng, and S. Yan · 2016
Later among the works it cites.
Emonets: Multimodal deep learning approaches for emotion recognition in video
S. E. Kahou, X. Bouthillier, P. Lamblin, C. Gulcehre, V. Michalski, K. Konda, S. Jean, P. Froumenty, Y. Dauphin, N. Boulanger-Lewandowski, et al · 2016
Later among the works it cites.
Moddrop: adaptive multi-modal gesture recognition
N. Neverova, C. Wolf, G. Taylor, and F. Nebout · 2016
Later among the works it cites.
Black holes and white rabbits: Metaphor identification with visual features
E. Shutova, D. Kiela, and J. Maillard · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Multimedia semantic integrity assessment using joint embedding of images and text
A. Jaiswal, E. Sabir, W. AbdAlmageed, and P. Natarajan · 2017
Later among the works it cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Later among the works it cites.
A batch learning framework for scalable personalized ranking
K. Liu and P. Natarajan · 2017
Later among the works it cites.
Multimodal named entity recognition for short social media posts
S. Moon, L. Neves, and V. Carvalho · 2018
Closest in time.