Fetching the paper…
Reading the bibliography…
Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
Bruno A. Olshausen and David J. Field · 1997
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2001
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2005
Earlier work this paper cites.
An information theoretic framework for multi-view learning
Karthik Sridharan and Sham M Kakade · 2008
Earlier work this paper cites.
Learning from multiple partially observed views-an application to multilingual text categorization
Massih R Amini, Nicolas Usunier, and Cyril Goutte · 2009
Earlier work this paper cites.
Linear spatial pyramid matching using sparse coding for image classification
Jianchao Yang, Kai Yu, Yihong Gong, and Thomas Huang · 2009
Earlier work this paper cites.
Online learning for matrix factorization and sparse coding, 2010
Julien Mairal, Francis Bach, Jean Ponce, and Guillermo Sapiro · 2010
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2012
Earlier work this paper cites.
Shift-invariance sparse coding for audio classification
Roger Grosse, Rajat Raina, Helen Kwong, and Andrew Y Ng · 2012
Earlier work this paper cites.
Visual classification with multitask joint sparse representation
Xiao-Tong Yuan, Xiaobai Liu, and Shuicheng Yan · 2012
Earlier work this paper cites.
A survey on multi-view learning
Chang Xu, Dacheng Tao, and Chao Xu · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Analyzing tensor power method dynamics in overcomplete regime, 2015
Anima Anandkumar, Rong Ge, and Majid Janzamin · 2015
Earlier work this paper cites.
Comparison and anti-concentration bounds for maxima of gaussian random vectors
Victor Chernozhukov, Denis Chetverikov, and Kengo Kato · 2015
Earlier work this paper cites.
Bounds on the expectation of the maximum of samples from a gaussian
Gautam Kamath · 2015
Earlier work this paper cites.
A multi-modal sparse coding classifier using dictionaries with different number of atoms
Soheil Shafiee, Farhad Kamangar, and Vassilis Athitsos · 2015
Earlier work this paper cites.
Auxiliary information regularized machine for multiple modality feature learning
Yang Yang, Han-Jia Ye, De-Chuan Zhan, and Yuan Jiang · 2015
Earlier work this paper cites.
Learning word representations with hierarchical sparse coding
Dani Yogatama, Manaal Faruqui, Chris Dyer, and Noah Smith · 2015
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2016
Cited alongside, same era.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Cited alongside, same era.
Multimodal sparse coding for event detection
Youngjune Gwon, William Campbell, Kevin Brady, Douglas Sturim, Miriam Cha, and HT Kung · 2016
Cited alongside, same era.
Heart sound classification via sparse coding
Bradley M Whitaker and David V Anderson · 2016
Cited alongside, same era.
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Later among the works it cites.
Cpm-nets: cross partial multi-view networks
Changqing Zhang, Zongbo Han, Yajie Cui, Huazhu Fu, Joey Tianyi Zhou, and Qinghua Hu · 2019
Later among the works it cites.
The sound of motions
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel · 2018
Cited alongside, same era.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Learning robust representations via multi-view information bottleneck
Marco Federici, Anjan Dutta, Patrick Forré, Nate Kushman, and Zeynep Akata · 2020
Later among the works it cites.
Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies, 2020
Itai Gat, Idan Schwartz, Alexander Schwing, and Tamir Hazan · 2020
Later among the works it cites.
Tcgm: An information-theoretic framework for semi-supervised multi-modality learning, 2020
Xinwei Sun, Yilun Xu, Peng Cao, Yuqing Kong, Lingjing Hu, Shanghang Zhang, and Yizhou Wang · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
Modality laziness: Everybody’s business is nobody’s business
Chenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu, Yue Wang, Yang Yuan, and Hang Zhao · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
What makes multi-modal learning better than single (provably), 2021
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang · 2021
Later among the works it cites.
M6: A chinese multimodal pretrainer
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et al · 2021
Later among the works it cites.
Towards understanding learning in neural networks with linear teachers
Roei Sarussi, Alon Brutzkus, and Amir Globerson · 2021
Later among the works it cites.
Toward understanding the feature learning process of self-supervised contrastive learning
Zixin Wen and Yuanzhi Li · 2021
Later among the works it cites.