Fetching the paper…
Reading the bibliography…
Multimodal learning methods with targeted unimodal learning objectives have exhibited their superior efficacy in alleviating the imbalanced multimodal learning problem.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Multiple-gradient descent algorithm (mgda) for multiobjective optimization
Désidéri, J.-A · 2012
Earlier work this paper cites.
Crema-d: Crowd-sourced emotional multimodal actors dataset
Cao, H., Cooper, D. G., Keutmann, M. K., Gur, R. C., Nenkova, A., and Verma, R · 2014
Earlier work this paper cites.
Multi-view convolutional neural networks for 3d shape recognition
Su, H., Maji, S., Kalogerakis, E., and Learned-Miller, E · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
Zadeh, A., Zellers, R., Pincus, E., and Morency, L.-P · 2016
Earlier work this paper cites.
Look, listen and learn
Arandjelovic, R. and Zisserman, A · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., McAllester, D., and Srebro, N · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sabour, S., Frosst, N., and Hinton, G. E · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Baltrušaitis, T., Ahuja, C., and Morency, L.-P · 2018
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Chen, Z., Badrinarayanan, V., Lee, C.-Y., and Rabinovich, A · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T · 2018
Cited alongside, same era.
Multi-task learning as multi-objective optimization
Sener, O. and Koltun, V · 2018
Cited alongside, same era.
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2018
Cited alongside, same era.
Learning not to learn: Training deep neural networks with biased data
Kim, B., Kim, H., Kim, K., Kim, S., and Kim, J · 2019
Cited alongside, same era.
Pareto multi-task learning
Modality competition: What makes joint training of multi-modal network fail in deep learning?(provably)
Huang, Y., Lin, J., Zhou, C., Yang, H., and Huang, L · 2022
Later among the works it cites.
Balanced multimodal learning via on-the-fly gradient modulation
Peng, X., Wei, Y., Deng, A., Wang, D., and Hu, D · 2022
Later among the works it cites.
Learning in audio-visual context: A review, analysis, and new perspective
Wei, Y., Hu, D., Tian, Y., and Li, X · 2022
Later among the works it cites.
Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks
Wu, N., Jastrzebski, S., Cho, K., and Geras, K. J · 2022
Later among the works it cites.
Vlp: A survey on vision-language pre-training
Chen, F.-L., Zhang, D.-Z., Han, M.-L., Chen, X.-Y., Shi, J., Xu, S., and Xu, B · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin, X., Zhen, H.-L., Li, Z., Zhang, Q.-F., and Kwong, S · 2019
Cited alongside, same era.
End-to-end multi-task learning with attention
Liu, S., Johns, E., and Davison, A. J · 2019
Cited alongside, same era.
Multimodal uncertainty reduction for intention recognition in human-robot interaction
Trick, S., Koert, D., Peters, J., and Rothkopf, C. A · 2019
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Cited alongside, same era.
Early-learning regularization prevents memorization of noisy labels
Liu, S., Niles-Weed, J., Razavian, N., and Fernandez-Granda, C · 2020
Cited alongside, same era.
Efficient continuous pareto exploration in multi-task learning
Ma, P., Du, T., and Matusik, W · 2020
Cited alongside, same era.
What makes training multi-modal classification networks hard?
Wang, W., Tran, D., and Feiszli, M · 2020
Cited alongside, same era.
On uni-modal feature learning in supervised multi-modal learning
Du, C., Teng, J., Li, T., Liu, Y., Yuan, T., Wang, Y., Yuan, Y., and Zhao, H · 2023
Later among the works it cites.
Pmr: Prototypical modal rebalance for multimodal learning
Fan, Y., Xu, W., Wang, H., Wang, J., and Guo, S · 2023
Later among the works it cites.
Boosting multi-modal model performance with adaptive gradient modulation
Li, H., Li, X., Hu, P., Lei, Y., Li, C., and Zhou, Y · 2023
Later among the works it cites.
Federated learning on multimodal data: A comprehensive survey
Lin, Y.-M., Gao, Y., Gong, M.-G., Zhang, S.-J., Zhang, Y.-Q., and Li, Z.-Y · 2023
Later among the works it cites.
Large-scale multi-modal pre-trained models: A comprehensive survey
Wang, X., Chen, G., Qian, G., Gao, P., Wei, X.-Y., Wang, Y., Tian, Y., and Gao, W · 2023
Later among the works it cites.
Mmcosine: Multi-modal cosine loss towards balanced audio-visual fine-grained learning
Xu, R., Feng, R., Zhang, S.-X., and Hu, D · 2023
Later among the works it cites.
Multimodal pretraining from monolingual to multilingual
Zhang, L., Ruan, L., Hu, A., and Jin, Q · 2023
Later among the works it cites.
Enhancing multimodal cooperation via sample-level modality valuation
Wei, Y., Feng, R., Wang, Z., and Hu, D · 2024
Closest in time.
Quantifying and enhancing multi-modal robustness with modality preference
Yang, Z., Wei, Y., Liang, C., and Hu, D · 2024
Closest in time.
Multimodal fusion on low-quality data: A comprehensive survey
Zhang, Q., Wei, Y., Han, Z., Fu, H., Peng, X., Deng, C., Hu, Q., Xu, C., Wen, J., Hu, D., et al · 2024
Closest in time.