Fetching the paper…
Reading the bibliography…
Multimodal learning helps to comprehensively understand the world, by integrating different senses.
Audio-visual automatic speech recognition: An overview
Gerasimos Potamianos, Chalapathy Neti, Juergen Luettin, and Iain Matthews · 2004
Earlier work this paper cites.
Cognitive neuroscience. the biology of the mind,(2014), 2006
Michael S Gazzaniga, Richard B Ivry, and GR Mangun · 2006
Earlier work this paper cites.
Deep boltzmann machines
Ruslan Salakhutdinov and Geoffrey Hinton · 2009
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Temporal multimodal learning in audiovisual speech recognition
Di Hu, Xuelong Li, et al · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Multimodal learning and reasoning for visual question answering
Ilija Ilievski and Jiashi Feng · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Earlier work this paper cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Later among the works it cites.
Listen to look: Action recognition by previewing audio
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani · 2020
Later among the works it cites.
Improving multimodal accuracy through modality pre-training and attention
Aya Abdelsalam Ismail, Mahmudul Hasan, and Faisal Ishtiaq · 2020
Later among the works it cites.
Speech2action: Cross-modal supervision for action recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, and Andrew Zisserman · 2020
Later among the works it cites.
x-vectors meet emotions: A study on dependencies between emotion and speaker recognition
Raghavendra Pappagari, Tianzi Wang, Jesus Villalba, Nanxin Chen, and Najim Dehak · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
Multimodal and temporal perception of audio-visual cues for emotion recognition
Esam Ghaleb, Mirela Popa, and Stylianos Asteriadis · 2019
Cited alongside, same era.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
Fengxiang He, Tongliang Liu, and Dacheng Tao · 2019
Cited alongside, same era.
A comprehensive study on bilingual and multilingual speech emotion recognition using a two-pass classification scheme
Panikos Heracleous and Akio Yoneyama · 2019
Cited alongside, same era.
Deep multimodal clustering for unsupervised audiovisual learning
Di Hu, Feiping Nie, and Xuelong Li · 2019
Cited alongside, same era.
Dense multimodal fusion for hierarchically joint representation
Di Hu, Chengze Wang, Feiping Nie, and Xuelong Li · 2019
Cited alongside, same era.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Cited alongside, same era.
Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classification and retrieval of videos
Kranti Parida, Neeraj Matiyali, Tanaya Guha, and Gaurav Sharma · 2020
Later among the works it cites.
Temporal relational modeling with self-supervision for action segmentation
Dong Wang, Di Hu, Xingjian Li, and Dejing Dou · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
On modality bias in the tvqa dataset
Thomas Winterbottom, Sarah Xiao, Alistair McLean, and Noura Al Moubayed · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as sgd
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer · 2020
Later among the works it cites.
Improving multi-modal learning with uni-modal teachers
Chenzhuang Du, Tingle Li, Yichen Liu, Zixin Wen, Tianyu Hua, Yue Wang, and Hang Zhao · 2021
Later among the works it cites.
Learning to balance the learning rates between various modalities via adaptive tracking factor
Ya Sun, Sijie Mai, and Haifeng Hu · 2021
Later among the works it cites.
Star: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Artificial neural variability for deep learning: On overfitting, noise memorization, and catastrophic forgetting
Zeke Xie, Fengxiang He, Shaopeng Fu, Issei Sato, Dacheng Tao, and Masashi Sugiyama · 2021
Later among the works it cites.
Positive sample propagation along the audio-visual event line
Jinxing Zhou, Liang Zheng, Yiran Zhong, Shijie Hao, and Meng Wang · 2021
Later among the works it cites.