Fetching the paper…
Reading the bibliography…
We abstract the features (i.e.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Smith, L. and Gasser, M · 2005
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
Learning from multiple partially observed views-an application to multilingual text categorization
Amini, M. R., Usunier, N., and Goutte, C · 2009
Earlier work this paper cites.
Multimodal deep learning
Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A. Y · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Silberman, N., Hoiem, D., Kohli, P., and Fergus, R · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
A survey on multi-view learning
Xu, C., Tao, D., and Xu, C · 2013
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Moddrop: adaptive multi-modal gesture recognition
Neverova, N., Wolf, C., Taylor, G., and Nebout, F · 2015
Earlier work this paper cites.
Stochastic optimization for multiview representation learning using partial least squares
Arora, R., Mianjy, P., and Marinov, T · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
Feichtenhofer, C., Pinz, A., and Zisserman, A · 2016
Earlier work this paper cites.
Cross modal distillation for supervision transfer
Gupta, S., Hoffman, J., and Malik, J · 2016
Earlier work this paper cites.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al · 2017
Cited alongside, same era.
Rdfnet: Rgb-d multi-level residual feature fusion for indoor semantic segmentation
Park, S.-J., Hong, K.-S., and Lee, S · 2017
Cited alongside, same era.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Agrawal, A., Batra, D., Parikh, D., and Kembhavi, A · 2018
Cited alongside, same era.
Modality distillation with multiple stream networks for action recognition
Garcia, N. C., Morerio, P., and Murino, V · 2018
Cited alongside, same era.
Graph distillation for action detection with privileged modalities
Luo, Z., Hsieh, J.-T., Jiang, L., Niebles, J. C., and Fei-Fei, L · 2018
Cited alongside, same era.
Efficient rgb-d semantic segmentation for indoor scene analysis
Seichter, D., Köhler, M., Lewandowski, B., Wengefeld, T., and Gross, H.-M · 2020
Later among the works it cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2020
Later among the works it cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
Tan, H. and Bansal, M · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Xiao, F., Lee, Y. J., Grauman, K., Malik, J., and Feichtenhofer, C · 2020
Later among the works it cites.
Knowledge distillation: A survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Acnet: Attention based network to exploit complementary features for rgbd semantic segmentation
Hu, X., Yang, K., Fei, L., and Wang, K · 2019
Cited alongside, same era.
Found in translation: Learning robust joint representations by cyclic translations between modalities
Pham, H., Liang, P. P., Manzini, T., Morency, L.-P., and Póczos, B · 2019
Cited alongside, same era.
Contrastive representation distillation
Tian, Y., Krishnan, D., and Isola, P · 2019
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2020
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Cited alongside, same era.
nnaudio: An on-the-fly gpu audio to spectrogram conversion toolbox using 1d convolutional neural networks
Cheuk, K. W., Anderson, H., Agres, K., and Herremans, D · 2020
Cited alongside, same era.
Large scale audiovisual learning of sounds with weakly labeled data
Fayek, H. M. and Kumar, A · 2020
Cited alongside, same era.
Later among the works it cites.
What makes multimodal learning better than single (provably)
Huang, Y., Du, C., Xue, Z., Chen, X., Zhao, H., and Huang, L · 2021
Later among the works it cites.
Multibench: Multiscale benchmarks for multimodal representation learning
Liang, P. P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M. A., Zhu, Y., et al · 2021
Later among the works it cites.
Attention bottlenecks for multimodal fusion
Nagrani, A., Yang, S., Arnab, A., Jansen, A., Schmid, C., and Sun, C · 2021
Later among the works it cites.
Adamml: Adaptive multi-modal learning for efficient video recognition
Panda, R., Chen, C.-F., Fan, Q., Sun, X., Saenko, K., Oliva, A., and Feris, R · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Huang, Y., Lin, J., Zhou, C., Yang, H., and Huang, L · 2022
Later among the works it cites.
Multiviz: An analysis benchmark for visualizing and understanding multimodal models
Liang, P. P., Lyu, Y., Chhablani, G., Jain, N., Deng, Z., Wang, X., Morency, L.-P., and Salakhutdinov, R · 2022
Later among the works it cites.
Balanced multimodal learning via on-the-fly gradient modulation
Peng, X., Wei, Y., Deng, A., Wang, D., and Hu, D · 2022
Later among the works it cites.
Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks
Wu, N., Jastrzebski, S., Cho, K., and Geras, K. J · 2022
Later among the works it cites.