Fetching the paper…
Reading the bibliography…
The complex world around us is inherently multimodal and sequential (continuous).
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 1906
Earlier work this paper cites.
Publicly available clinical BERT embeddings
Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott · 1909
Earlier work this paper cites.
Statistical inference for probabilistic functions of finite state markov chains
Leonard E Baum and Ted Petrie · 1966
Earlier work this paper cites.
Facial signs of emotional experience
Paul Ekman, Wallace V Freisen, and Sonia Ancoli · 1980
Earlier work this paper cites.
Applied regression analysis and other multivariable methods , volume 601
David G Kleinbaum, Lawrence L Kupper, Keith E Muller, and Azhar Nizam · 1988
Earlier work this paper cites.
An argument for basic emotions
Paul Ekman · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Early versus late fusion in semantic video analysis
Cees G. M. Snoek, Marcel Worring, and Arnold W. M. Smeulders · 2005
Earlier work this paper cites.
Latent-dynamic discriminative models for continuous gesture recognition
Louis-Philippe Morency, Ariadna Quattoni, and Trevor Darrell · 2007
Earlier work this paper cites.
Hidden conditional random fields
Ariadna Quattoni, Sybor Wang, Louis-Philippe Morency, Michael Collins, and Trevor Darrell · 2007
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette Chang, Sungbok Lee, and Shrikanth S. Narayanan · 2008
Earlier work this paper cites.
Speaker identification on the scotus corpus
Jiahong Yuan and Mark Liberman · 2008
Earlier work this paper cites.
Joint robust voicing detection and pitch estimation based on residual harmonics
Thomas Drugman and Abeer Alwan · 2011
Earlier work this paper cites.
Detection of glottal closure instants from speech signals: A quantitative review
Thomas Drugman, Mark Thomas, Jon Gudnason, Patrick Naylor, and Thierry Dutoit · 2012
Earlier work this paper cites.
Multi-view latent variable discriminative models for action recognition
Yale Song, Louis-Philippe Morency, and Randall Davis · 2012
Earlier work this paper cites.
Wavelet maxima dispersion for breathy to tense voice discrimination
John Kane and Christer Gobl · 2013
Earlier work this paper cites.
Action recognition by hierarchical sequence summarization
Yale Song, Louis-Philippe Morency, and Randall Davis · 2013
Cited alongside, same era.
A survey on multi-view learning
Chang Xu, Dacheng Tao, and Chao Xu · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Covarep—a collaborative voice analysis repository for speech technologies
Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Facial expression analysis, 2017
iMotions · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Temporal multimodal fusion for video emotion classification in the wild
Valentin Vielzeuf, Stéphane Pateux, and Frédéric Jurie · 2017
Later among the works it cites.
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2017
Later among the works it cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Computational analysis of persuasiveness in social multimedia: A novel dataset and multimodal prediction approach
Sunghyun Park, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, and Louis-Philippe Morency · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Cited alongside, same era.
Yang Gao, Oscar Beijbom, Ning Zhang, and Trevor Darrell · 2015
Cited alongside, same era.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, JungWoo Ha, and Byoung-Tak Zhang · 2016
Cited alongside, same era.
Deep multimodal fusion for persuasiveness prediction
Behnaz Nojavanasghari, Deepak Gopinath, Jayanth Koushik, Tadas Baltrušaitis, and Louis-Philippe Morency · 2016
Cited alongside, same era.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Multimodal language analysis with recurrent multistage fusion
Paul Pu Liang, Ziyin Liu, Amir Zadeh, and Louis-Philippe Morency · 2018
Later among the works it cites.
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2018
Later among the works it cites.
Multi-task learning of hierarchical vision-language representation
Duy-Kien Nguyen and Takayuki Okatani · 2018
Later among the works it cites.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, and Alexander Ku · 2018
Later among the works it cites.
Words can shift: Dynamically adjusting word representations using nonverbal behaviors
Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency · 2018
Later among the works it cites.
Multimodal language processing in human communication
Judith Holler and Stephen C Levinson · 2019
Closest in time.
CLEVR-dialog: A diagnostic dataset for multi-round reasoning in visual dialog
Satwik Kottur, José M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach · 2019
Closest in time.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Closest in time.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Closest in time.