Fetching the paper…
Reading the bibliography…
Intelligently reasoning about the world often requires integrating data from multiple modalities, as any individual modality may contain unreliable or incomplete information.
Convolutional networks for images, speech, and time-series
Yann Lecun and Yoshua Bengio · 1995
Earlier work this paper cites.
Multisensory contributions to low-level, ’unisensory’ processing
Charles E Schroeder and John Foxe · 2005
Earlier work this paper cites.
Multisensory processing via early cortical stages: Connections of the primary auditory cortical field with other sensory systems
E. Budinger, P. Heil, A. Hess, and H. Scheich · 2006
Earlier work this paper cites.
Multisensory processing in “unimodal” neurons: cross-modal subthreshold auditory effects in cat extrastriate visual cortex
Brian L Allman and M Alex Meredith · 2007
Earlier work this paper cites.
Subthreshold multisensory processing in cat auditory cortex
M Alex Meredith and Brian L Allman · 2009
Earlier work this paper cites.
Multimodal fusion for multimedia analysis: a survey
Pradeep K Atrey, M Anwar Hossain, Abdulmotaleb El Saddik, and Mohan S Kankanhalli · 2010
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Multimodal interaction: A review
Matthew Turk · 2014
Cited alongside, same era.
Audiovisual fusion: Challenges and new approaches
Aggelos K Katsaggelos, Sara Bahaadini, and Rafael Molina · 2015
Cited alongside, same era.
Recurrent neural networks for polyphonic sound event detection in real life recordings
G. Parascandolo, H. Huttunen, and T. Virtanen · 2016
Cited alongside, same era.
Audio-based multimedia event detection using deep recurrent neural networks
Y. Wang, L. Neves, and F. Metze · 2016
Cited alongside, same era.
Multimodal learning using 3d audio-visual data for audio-visual speech recognition
Rongfeng Su, Lan Wang, and Xunying Liu · 2017
Cited alongside, same era.
Jakobovski/free-spoken-digit-dataset: v1.0.8, August 2018
Zohar Jackson, César Souza, Jason Flaks, Yuxin Pan, Hereman Nicolas, and Adhish Thite · 2018
Later among the works it cites.
A survey of multi-view representation learning
Yingming Li, Ming Yang, and Zhongfei Zhang · 2018
Later among the works it cites.
Modality-invariant image-text embedding for image-sentence matching
Ruoyu Liu, Yao Zhao, Shikui Wei, Liang Zheng, and Yi Yang · 2019
Later among the works it cites.
A review of audio-visual fusion with machine learning
Xiaoyu Song, Hong Chen, Qing Wang, Yunqiang Chen, Mengxiao Tian, and Hui Tang · 2019
Later among the works it cites.
Modality attention for end-to-end audio-visual speech recognition
Pan Zhou, Wenwen Yang, Wei Chen, Yanfeng Wang, and Jia Jia · 2019
Later among the works it cites.
White noise analysis of neural networks
Ali Borji and Sikun Lin · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regions, periods, activities: Uncovering urban dynamics via cross-modal representation learning
Chao Zhang, Keyang Zhang, Quan Yuan, Haoruo Peng, Yu Zheng, Tim Hanratty, Shaowen Wang, and Jiawei Han · 2017
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Mmss: Multi-modal sharable and specific feature learning for rgb-d object recognition
Anran Wang, Jianfei Cai, Jiwen Lu, and Tat-Jen Cham
Cited in the paper.
Large-margin multi-modal deep learning for rgb-d object recognition
Anran Wang, Jiwen Lu, Jianfei Cai, Tat-Jen Cham, and Gang Wang
Cited in the paper.
Cheul Young Park, Narae Cha, Soowon Kang, Auk Kim, Ahsan Habib Khandoker, Leontios Hadjileontiadis, Alice Oh, Yong Jeong, and Uichin Lee · 2020
Closest in time.
Audio-visual recognition of overlapped speech for the lrs2 dataset
Jianwei Yu, Shi-Xiong Zhang, Jian Wu, Shahram Ghorbani, Bo Wu, Shiyin Kang, Shansong Liu, Xunying Liu, Helen Meng, and Dong Yu · 2020
Closest in time.