Fetching the paper…
Reading the bibliography…
Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Extraction of visual features for lipreading
Iain Matthews, Timothy F Cootes, J Andrew Bangham, Stephen Cox, and Richard Harvey. 2002 · 2002
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
Deep boltzmann machines
Ruslan Salakhutdinov and Geoffrey Hinton. 2009 · 2009
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. 2011 · 2011
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan R Salakhutdinov. 2012 · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng. 2013 · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Comparison of classifiers for lip reading with cuave and tulips database
Sunil S. Morade and Suprava Patnaik. 2015 · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Correlational neural networks
Sarath Chandar, Mitesh M Khapra, Hugo Larochelle, and Balaraman Ravindran. 2016 · 2016
Cited alongside, same era.
Multi30k: Multilingual english-german image descriptions
D. Elliott, S. Frank, K. Sima’an, and L. Specia. 2016 · 2016
Multi-lingual common semantic space construction via cluster-consistent word embedding
Lifu Huang, Kyunghyun Cho, Boliang Zhang, Heng Ji, and Kevin Knight. 2018 · 2018
Later among the works it cites.
On the implicit assumptions of gans
Ke Li and Jitendra Malik. 2018 · 2018
Later among the works it cites.
Efficient low-rank multimodal fusion with modality-specific factors
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency. 2018 · 2018
Later among the works it cites.
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze. 2018 · 2018
Later among the works it cites.
Non-parametric estimation of jensen-shannon divergence in generative adversarial network training
Mathieu Sinn and Ambrish Rawat. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
NIPS 2016 Tutorial: Generative adversarial networks
Ian Goodfellow. 2016 · 2016
Cited alongside, same era.
Generalization and equilibrium in generative adversarial nets (gans)
Sanjeev Arora, Rong Ge, Yingyu Liang, Tengyu Ma, and Yi Zhang. 2017 · 2017
Cited alongside, same era.
Findings of the Second Shared Task on Multimodal Machine Translation and Multilingual Image Description
Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Cuave: A new audio-visual database for multimodal human-computer interface research
Eric K Patterson, Sabri Gurbuz, Zekeriya Tufekci, and John N Gowdy. 2002 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017 · 2017
Cited alongside, same era.
Multimodal speech emotion recognition using audio and text
Seunghyun Yoon, Seokhyun Byun, and Kyomin Jung. 2018 · 2018
Later among the works it cites.
Jointly optimizing diversity and relevance in neural response generation
Xiang Gao, Sungjin Lee, Yizhe Zhang, Chris Brockett, Michel Galley, Jianfeng Gao, and Bill Dolan. 2019 · 2019
Closest in time.
Learning representations from imperfect time series data via tensor rank regularization
Paul Pu Liang, Zhun Liu, Yao-Hung Hubert Tsai, Qibin Zhao, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2019 · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Closest in time.
Found in translation: Learning robust joint representations by cyclic translations between modalities
Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabás Póczos. 2019 · 2019
Closest in time.
Variational mixture-of-experts autoencoders for multi-modal deep generative models
Yuge Shi, N Siddharth, Brooks Paige, and Philip Torr. 2019 · 2019
Closest in time.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Closest in time.
Speech emotion recognition using multi-hop attention mechanism
Seunghyun Yoon, Seokhyun Byun, Subhadeep Dey, and Kyomin Jung. 2019 · 2019
Closest in time.