Fetching the paper…
Reading the bibliography…
Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors.
Facial signs of emotional experience
Paul Ekman, Wallace V Freisen, and Sonia Ancoli. 1980 · 1980
Earlier work this paper cites.
An argument for basic emotions
Paul Ekman. 1992 · 1992
Earlier work this paper cites.
Tools, language and cognition in human evolution
Kathleen R Gibson, Kathleen Rita Gibson, and Tim Ingold. 1994 · 1994
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
Speaker identification on the scotus corpus
Jiahong Yuan and Mark Liberman. 2008 · 2008
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan R Salakhutdinov. 2012 · 2012
Earlier work this paper cites.
Covarepa collaborative voice analysis repository for speech technologies
Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer. 2014 · 2014
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Prismatic Inc, Steven J. Bethard, and David Mcclosky. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur P Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Online segment to segment neural transduction
Lei Yu, Jan Buys, and Phil Blunsom. 2016 · 2016
Cited alongside, same era.
Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016 · 2016
Cited alongside, same era.
Facial expression analysis
iMotions. 2017 · 2017
Cited alongside, same era.
Multimodal language analysis with recurrent multistage fusion
Paul Pu Liang, Ziyin Liu, Amir Zadeh, and Louis-Philippe Morency. 2018 · 2018
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Sentiment analysis using imperfect views from spoken language and acoustic modalities
Imran Sheikh, Sri Harsha Dumpala, Rupayan Chakraborty, and Sunil Kumar Kopparapu. 2018 · 2018
Later among the works it cites.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
Why self-attention? a targeted evaluation of neural machine translation architectures
Gongbo Tang, Mathias Müller, Annette Rios, and Rico Sennrich. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Context-dependent sentiment analysis in user-generated videos
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The best of both worlds: Combining recent advances in neural machine translation
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. 2018 · 2018
Cited alongside, same era.
Transformer-xl: Language modeling with longer-term dependency
Zihang Dai, Zhilin Yang, Yiming Yang, William W Cohen, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Multimodal affective analysis using hierarchical attention strategy with word-level alignment
Yue Gu, Kangning Yang, Shiyu Fu, Shuhong Chen, Xinyu Li, and Ivan Marsic. 2018 · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018 · 2018
Later among the works it cites.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli. 2019 · 2019
Closest in time.
Audio-visual fusion for sentiment classification using cross-modal autoencoder
Sri Harsha Dumpala, Rupayan Chakraborty, and Sunil Kumar Kopparapu. 2019 · 2019
Closest in time.
Found in translation: Learning robust joint representations by cyclic translations between modalities
Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabas Poczos. 2019 · 2019
Closest in time.
Learning factorized multimodal representations
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Closest in time.
Words can shift: Dynamically adjusting word representations using nonverbal behaviors
Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2019 · 2019
Closest in time.