Fetching the paper…
Reading the bibliography…
This paper explores a novel multi-modal alternating learning paradigm pursuing a reconciliation between the exploitation of uni-modal features and the exploration of cross-modal interactions.
On information and sufficiency
Kullback, S. and Leibler, R. A · 1951
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Boosting a weak learning algorithm by majority
Freund, Y · 1995
Earlier work this paper cites.
Experiments with a new boosting algorithm
Freund, Y., Schapire, R. E., et al · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Freund, Y. and Schapire, R. E · 1997
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Friedman, J. H · 2001
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Heterogeneous image feature integration via multi-modal spectral clustering
Cai, X., Nie, F., Huang, H., and Kamangar, F · 2011
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Cao, H., Cooper, D. G., Keutmann, M. K., Gur, R. C., Nenkova, A., and Verma, R · 2014
Earlier work this paper cites.
Covarep — a collaborative voice analysis repository for speech technologies
Degottex, G., Kane, J., Drugman, T., Raitio, T., and Scherer, S · 2014
Earlier work this paper cites.
Selfieboost: A boosting algorithm for deep learning
Shalev-Shwartz, S · 2014
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., and Nieto, O · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
Zadeh, A., Zellers, R., Pincus, E., and Morency, L.-P · 2016
Earlier work this paper cites.
Learning discrete representations via information maximizing self-augmented training
Hu, W., Miyato, T., Tokui, S., Matsumoto, E., and Sugiyama, M · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Earlier work this paper cites.
Deep multimodal feature analysis for action recognition in rgb+ d videos
Shahroudy, A., Ng, T.-T., Gong, Y., and Wang, G · 2017
Earlier work this paper cites.
Tri-clustered tensor completion for social-aware image tag refinement
Tang, J., Shu, X., Qi, G.-J., Li, Z., Wang, M., Yan, S., and Jain, R · 2017
Earlier work this paper cites.
Openface 2.0: Facial behavior analysis toolkit
Baltrusaitis, T., Zadeh, A., Lim, Y. C., and Morency, L.-P · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Guan, D., Cao, Y., Liang, J., Cao, Y., and Yang, M. Y · 2018
Cited alongside, same era.
Learning deep resnet blocks sequentially using boosting theory
Huang, F., Ash, J., Langford, J., and Schapire, R · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos
Tian, Y., Shi, J., Li, B., Duan, Z., and Xu, C · 2018
Cited alongside, same era.
Recognizing emotions in video using multimodal dnn feature fusion
Williams, J., Kleinegesse, S., Comanescu, R., and Radu, O · 2018
Exploring complex and heterogeneous correlations on hypergraph for the prediction of drug-target interactions
Ruan, D., Ji, S., Yan, C., Zhu, J., Zhao, X., Yang, Y., Gao, Y., Zou, C., and Dai, Q · 2021
Later among the works it cites.
Efficient rgb-d semantic segmentation for indoor scene analysis
Seichter, D., Köhler, M., Lewandowski, B., Wengefeld, T., and Gross, H.-M · 2021
Later among the works it cites.
Pan-cancer integrative histology-genomic analysis via multimodal deep learning
Chen, R. J., Lu, M. Y., Williamson, D. F., Chen, T. Y., Lipkova, J., Noor, Z., Shaban, M., Shady, M., Williams, M., Joo, B., et al · 2022
Later among the works it cites.
Shrec’22 track: Open-set 3d object retrieval
Feng, Y., Gao, Y., Zhao, X., Guo, Y., Bagewadi, N., Bui, N.-T., Dao, H., Gangisetty, S., Guan, R., Han, X., et al · 2022
Later among the works it cites.
Event-based vision: A survey
Gallego, G., Delbruck, T., Orchard, G., Bartolozzi, C., Taba, B., Censi, A., Leutenegger, S., Davison, A. J., Conradt, J., Daniilidis, K., and Scaramuzza, D · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
Zadeh, A., Liang, P. P., Poria, S., Cambria, E., and Morency, L.-P · 2018
Cited alongside, same era.
Multimodal machine learning: A survey and taxonomy
Baltrušaitis, T., Ahuja, C., and Morency, L.-P · 2019
Cited alongside, same era.
Dm2c: Deep mixed-modal clustering
Jiang, Y., Xu, Q., Yang, Z., Cao, X., and Huang, Q · 2019
Cited alongside, same era.
Deep collaborative embedding for social image understanding
Li, Z., Tang, J., and Mei, T · 2019
Cited alongside, same era.
Adagcn: Adaboosting graph convolutional networks into deep models
Sun, K., Zhu, Z., and Lin, Z · 2019
Cited alongside, same era.
Gradient boosting neural networks: Grownet
Badirli, S., Liu, X., Xing, Z., Bhowmik, A., and Keerthi, S. S · 2020
Cited alongside, same era.
Bi-directional cross-modality feature propagation with separation-and-aggregation gate for rgb-d semantic segmentation
Chen, X., Lin, K.-Y., Wang, J., Wu, W., Qian, C., Li, H., and Zeng, G · 2020
Cited alongside, same era.
Modality competition: What makes joint training of multi-modal network fail in deep learning? (Provably)
Huang, Y., Lin, J., Zhou, C., Yang, H., and Huang, L · 2022
Later among the works it cites.
Modeling multiple views via implicitly preserving global consistency and local complementarity
Li, J., Qiang, W., Zheng, C., Su, B., Razzak, F., Wen, J.-R., and Xiong, H · 2022
Later among the works it cites.
Balanced multimodal learning via on-the-fly gradient modulation
Peng, X., Wei, Y., Deng, A., Wang, D., and Hu, D · 2022
Later among the works it cites.
Learning in audio-visual context: A review, analysis, and new perspective, 2022
Wei, Y., Hu, D., Tian, Y., and Li, X · 2022
Later among the works it cites.
Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks
Wu, N., Jastrzebski, S., Cho, K., and Geras, K. J · 2022
Later among the works it cites.
Avqa: A dataset for audio-visual question answering on videos
Yang, P., Wang, X., Duan, X., Chen, H., Hou, R., Jin, C., and Zhu, W · 2022
Later among the works it cites.
On uni-modal feature learning in supervised multi-modal learning
Du, C., Teng, J., Li, T., Liu, Y., Yuan, T., Wang, Y., Yuan, Y., and Zhao, H · 2023
Later among the works it cites.
Pmr: Prototypical modal rebalance for multimodal learning
Fan, Y., Xu, W., Wang, H., Wang, J., and Guo, S · 2023
Later among the works it cites.
Superfast: 200× video frame interpolation via event camera
Gao, Y., Li, S., Li, Y., Guo, Y., and Dai, Q · 2023
Later among the works it cites.
Identity-invariant representation and transformer-style relation for micro-expression recognition
Shao, Z., Li, F., Zhou, Y., Chen, H., Zhu, H., and Yao, R · 2023
Later among the works it cites.
Rpeflow: Multimodal fusion of rgb-pointcloud-event for joint optical flow and scene flow estimation
Wan, Z., Mao, Y., Zhang, J., and Dai, Y · 2023
Later among the works it cites.
Provable dynamic fusion for low-quality multimodal data
Zhang, Q., Wu, H., Zhang, C., Hu, Q., Fu, H., Zhou, J. T., and Peng, X · 2023
Later among the works it cites.
Joint facial action unit recognition and self-supervised optical flow estimation
Shao, Z., Zhou, Y., Li, F., Zhu, H., and Liu, B · 2024
Closest in time.
Rnve: A real nighttime vision enhancement benchmark and dual-stream fusion network
Wang, Y., Zhang, Y., Guo, Q., Zhao, M., and Jiang, Y · 2024
Closest in time.
Multimodal fusion on low-quality data: A comprehensive survey, 2024
Zhang, Q., Wei, Y., Han, Z., Fu, H., Peng, X., Deng, C., Hu, Q., Xu, C., Wen, J., Hu, D., and Zhang, C · 2024
Closest in time.