Fetching the paper…
Reading the bibliography…
The world provides us with data of multiple modalities.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Linda Smith and Michael Gasser · 2005
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan · 2008
Earlier work this paper cites.
Learning from multiple sources
Koby Crammer, Michael Kearns, and Jennifer Wortman · 2008
Earlier work this paper cites.
An information theoretic framework for multi-view learning
Karthik Sridharan and Sham M Kakade · 2008
Earlier work this paper cites.
Learning from multiple partially observed views-an application to multilingual text categorization
Massih-Reza Amini, Nicolas Usunier, Cyril Goutte, et al · 2009
Earlier work this paper cites.
Linear algorithms for online multitask classification
Giovanni Cavallanti, Nicolo Cesa-Bianchi, and Claudio Gentile · 2010
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Linear regression analysis
George AF Seber and Alan J Lee · 2012
Earlier work this paper cites.
Excess risk bounds for multitask learning with trace norm regularization
Massimiliano Pontil and Andreas Maurer · 2013
Earlier work this paper cites.
A survey on multi-view learning
Chang Xu, Dacheng Tao, and Chao Xu · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Auxiliary information regularized machine for multiple modality feature learning
Yang Yang, Han-Jia Ye, De-Chuan Zhan, and Yuan Jiang · 2015
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman · 2016
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Saurabh Gupta, Judy Hoffman, and Jitendra Malik · 2016
Earlier work this paper cites.
Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture
Caner Hazirbas, Lingni Ma, Csaba Domokos, and Daniel Cremers · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J Foster, and Matus Telgarsky · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Cpm-nets: Cross partial multi-view networks
Changqing Zhang, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu, et al · 2019
Later among the works it cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Later among the works it cites.
Few-shot learning via learning the representation, provably
Simon S Du, Wei Hu, Sham M Kakade, Jason D Lee, and Qi Lei · 2020
Later among the works it cites.
Large scale audiovisual learning of sounds with weakly labeled data
Haytham M Fayek and Anurag Kumar · 2020
Later among the works it cites.
Learning robust representations via multi-view information bottleneck
Marco Federici, Anjan Dutta, Patrick Forré, Nate Kushman, and Zeynep Akata · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rdfnet: Rgb-d multi-level residual feature fusion for indoor semantic segmentation
Seong-Jin Park, Ki-Sang Hong, and Seungyong Lee · 2017
Cited alongside, same era.
Context-dependent sentiment analysis in user-generated videos
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency · 2017
Cited alongside, same era.
Learning cross-modal deep representations for robust pedestrian detection
Dan Xu, Wanli Ouyang, Elisa Ricci, Xiaogang Wang, and Nicu Sebe · 2017
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel · 2018
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Rednet: Residual encoder-decoder network for indoor rgb-d semantic segmentation
Jindong Jiang, Lunan Zheng, Fei Luo, and Zhijun Zhang · 2018
Cited alongside, same era.
Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies
Itai Gat, Idan Schwartz, Alexander Schwing, and Tamir Hazan · 2020
Later among the works it cites.
Late temporal modeling in 3d cnn architectures with bert for action recognition
M Esat Kalfaoglu, Sinan Kalkan, and A Aydin Alatan · 2020
Later among the works it cites.
Efficient rgb-d semantic segmentation for indoor scene analysis
Daniel Seichter, Mona Köhler, Benjamin Lewandowski, Tim Wengefeld, and Horst-Michael Gross · 2020
Later among the works it cites.
Tcgm: An information-theoretic framework for semi-supervised multi-modality learning
Xinwei Sun, Yilun Xu, Peng Cao, Yuqing Kong, Lingjing Hu, Shanghang Zhang, and Yizhou Wang · 2020
Later among the works it cites.
Provable meta-learning of linear representations
Nilesh Tripuraneni, Chi Jin, and Michael I Jordan · 2020
Later among the works it cites.
On the theory of transfer learning: The importance of task diversity
Nilesh Tripuraneni, Michael Jordan, and Chi Jin · 2020
Later among the works it cites.
Self-supervised learning from a multi-view perspective
Yao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Deep multimodal fusion by channel exchanging
Yikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu, Yu Rong, and Junzhou Huang · 2020
Later among the works it cites.
Tao Zhou, Kim-Han Thung, Mingxia Liu, Feng Shi, Changqing Zhang, and Dinggang Shen · 2020
Later among the works it cites.
Improving multi-modal learning with uni-modal teachers
Chenzhuang Du, Tingle Li, Yichen Liu, Zixin Wen, Tianyu Hua, Yue Wang, and Hang Zhao · 2021
Closest in time.
Trusted multi-view classification
Zongbo Han, Changqing Zhang, Huazhu Fu, and Joey Tianyi Zhou · 2021
Closest in time.