Fetching the paper…
Reading the bibliography…
Computational modeling of human multimodal language is an emerging research area in natural language processing spanning the language, visual and acoustic modalities.
Syntactic Structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Statistical inference for probabilistic functions of finite state markov chains
Leonard E Baum and Ted Petrie. 1966 · 1966
Earlier work this paper cites.
Joint robust voicing detection and pitch estimation based on residual harmonics
Thomas Drugman and Abeer Alwan. 2011 · 1976
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Vocal quality factors: Analysis, synthesis, and perception
Donald G Childers and CK Lee. 1991 · 1991
Earlier work this paper cites.
Glottal wave analysis with pitch synchronous iterative adaptive inverse filtering
Paavo Alku. 1992 · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
Parabolic spectral parameterâa new method for quantification of the glottal flow
Paavo Alku, Helmer Strik, and Erkki Vilkman. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
M. Schuster and K.K. Paliwal. 1997 · 1997
Earlier work this paper cites.
Recurrent Neural Networks: Design and Applications , 1st edition
L. C. Jain and L. R. Medsker. 1999 · 1999
Earlier work this paper cites.
A new view of language acquisition
Patricia K. Kuhl. 2000 · 2000
Earlier work this paper cites.
Random forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Normalized amplitude quotient for parametrization of the glottal flow
Paavo Alku, Tom Bäckström, and Erkki Vilkman. 2002 · 2002
Earlier work this paper cites.
Fast human detection using a cascade of histograms of oriented gradients
Qiang Zhu, Mei-Chen Yeh, Kwang-Ting Cheng, and Shai Avidan. 2006 · 2006
Earlier work this paper cites.
Trueskill™: A bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. 2007 · 2007
Earlier work this paper cites.
Latent-dynamic discriminative models for continuous gesture recognition
Louis-Philippe Morency, Ariadna Quattoni, and Trevor Darrell. 2007 · 2007
Earlier work this paper cites.
Hidden conditional random fields
Ariadna Quattoni, Sybor Wang, Louis-Philippe Morency, Michael Collins, and Trevor Darrell. 2007 · 2007
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008 · 2008
Earlier work this paper cites.
Speaker identification on the scotus corpus
Jiahong Yuan and Mark Liberman. 2008 · 2008
Earlier work this paper cites.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Towards multimodal sentiment analysis: Harvesting opinions from the web
Louis-Philippe Morency, Rada Mihalcea, and Payal Doshi. 2011 · 2011
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011 · 2011
Earlier work this paper cites.
Detection of glottal closure instants from speech signals: A quantitative review
Thomas Drugman, Mark Thomas, Jon Gudnason, Patrick Naylor, and Thierry Dutoit. 2012 · 2012
Earlier work this paper cites.
Multi-view latent variable discriminative models for action recognition
Yale Song, Louis-Philippe Morency, and Randall Davis. 2012 · 2012
Cited alongside, same era.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan R Salakhutdinov. 2012 · 2012
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A. r. Mohamed, and G. Hinton. 2013 · 2013
Cited alongside, same era.
Wavelet maxima dispersion for breathy to tense voice discrimination
John Kane and Christer Gobl. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Action recognition by hierarchical sequence summarization
Yale Song, Louis-Philippe Morency, and Randall Davis. 2013 · 2013
End-to-end sequence labeling via bi-directional lstm-cnns-crf
Xuezhe Ma and Eduard Hovy. 2016 · 2016
Later among the works it cites.
Deep multimodal fusion for persuasiveness prediction
Behnaz Nojavanasghari, Deepak Gopinath, Jayanth Koushik, Tadas Baltrušaitis, and Louis-Philippe Morency. 2016 · 2016
Later among the works it cites.
Convolutional mkl based multimodal emotion recognition and sentiment analysis
Soujanya Poria, Iti Chaturvedi, Erik Cambria, and Amir Hussain. 2016 · 2016
Later among the works it cites.
Extending long short-term memory for multi-view structured learning
Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrušaitis, and Goecke Roland. 2016 · 2016
Later among the works it cites.
Neural speech recognizer: Acoustic-to-word LSTM model for large vocabulary speech recognition
Hagen Soltau, Hank Liao, and Hasim Sak. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey on multi-view learning
Chang Xu, Dacheng Tao, and Chao Xu. 2013 · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Covarepâa collaborative voice analysis repository for speech technologies
Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer. 2014 · 2014
Cited alongside, same era.
Speech emotion recognition using deep neural network and extreme learning machine
Kun Han, Dong Yu, and Ivan Tashev. 2014 · 2014
Cited alongside, same era.
Computational analysis of persuasiveness in social multimedia: A novel dataset and multimodal prediction approach
Sunghyun Park, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, and Louis-Philippe Morency. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Cited alongside, same era.
Look, listen, and decode: Multimodal speech recognition with images
F. Sun, D. Harwath, and J. Glass. 2016 · 2016
Later among the works it cites.
Select-additive learning: Improving cross-individual generalization in multimodal sentiment analysis
Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, and Eric P Xing. 2016 · 2016
Later among the works it cites.
Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016 · 2016
Later among the works it cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber. 2016 · 2016
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2017 · 2017
Later among the works it cites.
Multimodal sentiment analysis with word-level fusion and reinforcement learning
Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, and Louis-Philippe Morency. 2017 · 2017
Later among the works it cites.
Visual features for context-aware speech recognition
Abhinav Gupta, Yajie Miao, Leonardo Neves, and Florian Metze. 2017 · 2017
Later among the works it cites.
Learning word-like units from joint audio-visual analysis
David F. Harwath and James R. Glass. 2017 · 2017
Later among the works it cites.
Facial expression analysis
iMotions. 2017 · 2017
Later among the works it cites.
Visually grounded learning of keyword prediction from untranscribed speech
Herman Kamper, Shane Settle, Gregory Shakhnarovich, and Karen Livescu. 2017 · 2017
Later among the works it cites.
Character-based bidirectional lstm-crf with words and characters for japanese named entity recognition
Shotaro Misawa, Motoki Taniguchi, Yasuhide Miura, and Tomoko Ohkuma. 2017 · 2017
Later among the works it cites.
Emergence of multimodal action representations from neural network self-organization
German I. Parisi, Jun Tani, Cornelius Weber, and Stefan Wermter. 2017 · 2017
Later among the works it cites.
Context-dependent sentiment analysis in user-generated videos
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017 · 2017
Later among the works it cites.
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017 · 2017
Later among the works it cites.
Reusing neural speech representations for auditory emotion recognition
Egor Lakomkin, Cornelius Weber, Sven Magg, and Stefan Wermter. 2018 · 2018
Closest in time.
Multimodal local-global ranking fusion for emotion recognition
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2018 · 2018
Closest in time.
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency. 2018 · 2018
Closest in time.
Seq2seq2sentiment: Multimodal sequence to sequence models for sentiment analysis
Hai Pham, Thomas Manzini, Paul Pu Liang, and Barnabas Poczos. 2018 · 2018
Closest in time.
Learning factorized multimodal representations
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2018 · 2018
Closest in time.