Fetching the paper…
Reading the bibliography…
Standard multi-modal models assume the use of the same modalities in training and inference stages.
Dmcl: Distillation multiple choice learning for multimodal action recognition
Garcia, N. C.; Bargal, S. A.; Ablavsky, V.; Morerio, P.; Murino, V.; and Sclaroff, S. 2019 · 1912
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
A new learning paradigm: Learning using privileged information
Vapnik, V.; and Vashist, A. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
On the theory of learnining with privileged information
Pechyony, D.; and Vapnik, V. 2010 · 2010
Earlier work this paper cites.
Cross-view action modeling, learning and recognition
Wang, J.; Nie, X.; Xia, Y.; Wu, Y.; and Zhu, S.-C. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; Dean, J.; et al. 2015 · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
Lopez-Paz, D.; Bottou, L.; Schölkopf, B.; and Vapnik, V. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; and Paluri, M. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Earlier work this paper cites.
Learning with side information through modality hallucination
Hoffman, J.; Gupta, S.; and Darrell, T. 2016 · 2016
Earlier work this paper cites.
Histogram of oriented principal components for cross-view action recognition
Rahmani, H.; Mahmood, A.; Huynh, D.; and Mian, A. 2016 · 2016
Earlier work this paper cites.
Ntu rgb+ d: A large scale dataset for 3d human activity analysis
Shahroudy, A.; Liu, J.; Ng, T.-T.; and Wang, G. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; and Wojna, Z. 2016 · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Gool, L. V. 2016 · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Cited alongside, same era.
X3d: Expanding architectures for efficient video recognition
Feichtenhofer, C. 2020 · 2020
Later among the works it cites.
D3d: Distilled 3d networks for video action recognition
Stroud, J.; Ross, D.; Sun, C.; Deng, J.; and Sukthankar, R. 2020 · 2020
Later among the works it cites.
Vivit: A video vision transformer
Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lučić, M.; and Schmid, C. 2021 · 2021
Later among the works it cites.
Multimodal fusion via teacher-student network for indoor action recognition
Bruce, X.; Liu, Y.; and Chan, K. C. 2021 · 2021
Later among the works it cites.
What makes multi-modal learning better than single (provably)
Huang, Y.; Du, C.; Xue, Z.; Chen, X.; Zhao, H.; and Huang, L. 2021 · 2021
Later among the works it cites.
SMIL: Multimodal learning with severely missing modality
Ma, M.; Ren, J.; Zhao, L.; Tulyakov, S.; Wu, C.; and Peng, X. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Modality distillation with multiple stream networks for action recognition
Garcia, N. C.; Morerio, P.; and Murino, V. 2018 · 2018
Cited alongside, same era.
Deep bilinear learning for rgb-d action recognition
Hu, J.-F.; Zheng, W.-S.; Pan, J.; Lai, J.; and Zhang, J. 2018 · 2018
Cited alongside, same era.
Graph distillation for action detection with privileged modalities
Luo, Z.; Hsieh, J.-T.; Jiang, L.; Niebles, J. C.; and Fei-Fei, L. 2018 · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Tran, D.; Wang, H.; Torresani, L.; Ray, J.; LeCun, Y.; and Paluri, M. 2018 · 2018
Cited alongside, same era.
Non-local neural networks
Wang, X.; Girshick, R.; Gupta, A.; and He, K. 2018 · 2018
Cited alongside, same era.
Slowfast networks for video recognition
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019 · 2019
Cited alongside, same era.
Learning with privileged information via adversarial discriminative modality distillation
Garcia, N. C.; Morerio, P.; and Murino, V. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Multi-modality learning for human action recognition
Ren, Z.; Zhang, Q.; Gao, X.; Hao, P.; and Cheng, J. 2021 · 2021
Later among the works it cites.
Missing modality imagination network for emotion recognition with uncertain missing modalities
Zhao, J.; Li, R.; and Jin, Q. 2021 · 2021
Later among the works it cites.
MultiMAE: Multi-modal Multi-task Masked Autoencoders
Bachmann, R.; Mizrahi, D.; Atanov, A.; and Zamir, A. 2022 · 2022
Closest in time.
MMNet: A Model-based Multimodal Network for Human Action Recognition in RGB-D Videos
Bruce, X.; Liu, Y.; Zhang, X.; Zhong, S.-h.; and Chan, K. C. 2022 · 2022
Closest in time.
Masked Autoencoders As Spatiotemporal Learners
Feichtenhofer, C.; Fan, H.; Li, Y.; and He, K. 2022 · 2022
Closest in time.
Masked autoencoders are scalable vision learners
He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; and Girshick, R. 2022 · 2022
Closest in time.
Are Multimodal Transformers Robust to Missing Modality?
Ma, M.; Ren, J.; Zhao, L.; Testuggine, D.; and Peng, X. 2022 · 2022
Closest in time.
3dfcnn: Real-time action recognition using 3d deep neural networks with raw depth information
Sanchez-Caballero, A.; de López-Diz, S.; Fuentes-Jimenez, D.; Losada-Gutiérrez, C.; Marrón-Romera, M.; Casillas-Perez, D.; and Sarker, M. I. 2022 · 2022
Closest in time.
Learning using privileged information: similarity control and knowledge transfer
Vapnik, V.; Izmailov, R.; et al. 2015 · 2049
Closest in time.