Fetching the paper…
Reading the bibliography…
We present Masked Audio-Video Learners (MAViL) to train audio-visual representations.
T. Chen and R. R. Rao, “Audio-visual integration in multimodal communication,” Proceedings of the IEEE , vol. 86, no. 5, pp. 837–852, 1998
1998
Earlier work this paper cites.
D. Roy, “Learning from sights and sounds: a computational model,” PhD Thesis, MIT Media Laboratory , 1999
1999
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audiovisual speech,” Proceedings of the IEEE , vol. 91, no. 9, pp. 1306–1326, 2003
2003
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2005), 20-26 June 2005, San Diego, CA, USA . IEEE Computer Society, 2005, pp. 886–893
2005
Earlier work this paper cites.
P. S. Aleksic and A. K. Katsaggelos, “Audio-visual biometrics,” Proceedings of the IEEE , vol. 94, no. 11, pp. 2025–2044, 2006
2006
Earlier work this paper cites.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2010, Chia Laguna Resort, Sardinia, Italy, May 13-15, 2010 , ser. JMLR Proceedings, vol. 9. JMLR.org, 2010, pp. 297–304
2010
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in ICML , 2011
2011
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. CONF. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
Y. Kim, H. Lee, and E. M. Provost, “Deep learning for robust feature generation in audiovisual emotion recognition,” in ICASSP , 2013
2013
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” NeurIPS Deep Learning and Representation Learning Workshop , 2015
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for Environmental Sound Classification,” in Proceedings of the 23rd Annual ACM Conference on Multimedia . ACM Press, 2015, pp. 1015–1018
2015
Earlier work this paper cites.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in NeurIPS , 2016
2016
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and R. P. Wildes, “Spatiotemporal residual networks for video action recognition,” in NIPS , 2016
2016
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “MSR-VTT: A large video description dataset for bridging video and language,” in CVPR , 2016
2016
Earlier work this paper cites.
D. Ramachandram and G. W. Taylor, “Deep multimodal learning: A survey on recent advances and trends,” IEEE signal processing magazine , vol. 34, no. 6, pp. 96–108, 2017
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in ICCV , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. Weiss, and K. Wilson, “CNN architectures for large-scale audio classification,” in ICASSP , 2017
2017
Earlier work this paper cites.
——, “SGDR: Stochastic gradient descent with warm restarts,” in ICLR , 2017
2017
Earlier work this paper cites.
G. Larsson, M. Maire, and G. Shakhnarovich, “FractalNet: Ultra-deep neural networks without residuals,” in ICLR , 2017
2017
Earlier work this paper cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in ECCV , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, “Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,” ACM Transactions on Graphics (TOG) , vol. 37, no. 4, pp. 1–11, 2018
2018
Earlier work this paper cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,” in ECCV , 2018
2018
Earlier work this paper cites.
——, “Objects that sound,” in ECCV , 2018
2018
Earlier work this paper cites.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Mishra and D. Marr, “Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,” in ICLR , 2018
2018
Earlier work this paper cites.
L. Zhou, C. Xu, and J. J. Corso, “Towards automatic learning of procedures from web instructional videos,” in AAAI , 2018
2018
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in ICLR , 2018
2018
Earlier work this paper cites.
P. Warden, “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition,” ArXiv e-prints , Apr. 2018
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Epic-fusion: Audio-visual temporal binding for egocentric action recognition,” in ICCV , 2019
2019
Earlier work this paper cites.
Y. Tian, D. Krishnan, and P. Isola, “Contrastive representation distillation,” in ICLR , 2019
2019
Cited alongside, same era.
W. Park, D. Kim, Y. Lu, and M. Cho, “Relational knowledge distillation,” in CVPR , 2019
2019
Cited alongside, same era.
J. H. Cho and B. Hariharan, “On the efficacy of knowledge distillation,” in ICCV , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,” in ICCV , 2019
2019
Cited alongside, same era.
P. Morgado, I. Misra, and N. Vasconcelos, “Robust audio-visual instance discrimination,” in CVPR , 2021
2021
Later among the works it cites.
X. Chen, S. Xie, and K. He, “An empirical study of training self-supervised vision transformers,” in ICCV , 2021
2021
Later among the works it cites.
S. Hershey, D. P. Ellis, E. Fonseca, A. Jansen, C. Liu, R. C. Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” in ICASSP , 2021
2021
Later among the works it cites.
Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio Spectrogram Transformer,” in Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in NAACL-HLT , 2019
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Cited alongside, same era.
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo, “Cutmix: Regularization strategy to train strong classifiers with localizable features,” in ICCV , 2019
2019
Cited alongside, same era.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in ICCV , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
S. Ma, Z. Zeng, D. McDuff, and Y. Song, “Active contrastive learning of audio-visual video representations,” in ICLR , 2020
2020
Cited alongside, same era.
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “Slow-fast auditory streams for audio recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 855–859
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira, “Perceiver: General perception with iterative attention,” in ICML , 2021
2021
Later among the works it cites.
A. Rouditchenko, A. Boggust, D. Harwath, B. Chen, D. Joshi, S. Thomas, K. Audhkhasi, H. Kuehne, R. Panda, R. Feris et al. , “AVLnet: Learning audio-visual language representations from instructional videos,” in Interspeech , 2021
2021
Later among the works it cites.
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Feichtenhofer, “Multiscale vision transformers,” in ICCV , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Oncescu, A. S. Koepke, J. F. Henriques, Z. Akata, and S. Albanie, “Audio retrieval with natural language queries,” in Interspeech 2021, 22nd Annual Conference of the International Speech Communication Association, Brno, Czechia, 30 August - 3 September 2021 . ISCA, 2021, pp. 2411–2415
2021
Later among the works it cites.
X. Mei, X. Liu, Q. Huang, M. D. Plumbley, and W. Wang, “Audio captioning transformer,” in Proceedings of the 6th Detection and Classification of Acoustic Scenes and Events 2021 Workshop (DCASE2021) , Barcelona, Spain, November 2021, pp. 211–215
2021
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR , 2022
2022
Closest in time.
C. Feichtenhofer, H. Fan, Y. Li, and K. He, “Masked autoencoders as spatiotemporal learners,” in NeurIPS , 2022
2022
Closest in time.
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, C. Feichtenhofer et al. , “Masked autoencoders that listen,” in NeurIPS , 2022
2022
Closest in time.
2022
Closest in time.
C. Wei, H. Fan, S. Xie, C.-Y. Wu, A. Yuille, and C. Feichtenhofer, “Masked feature prediction for self-supervised visual pre-training,” in CVPR , 2022
2022
Closest in time.
G. Chrupała, “Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques,” Journal of Artificial Intelligence Research , vol. 73, pp. 673–707, 2022
2022
Closest in time.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” in ICLR , 2022
2022
Closest in time.
2022
Closest in time.
L. Wei, L. Xie, W. Zhou, H. Li, and Q. Tian, “MVP: Multimodality-guided visual pre-training,” in ECCV , 2022
2022
Closest in time.
2022
Closest in time.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “Data2vec: A general framework for self-supervised learning in speech, vision and language,” in ICML , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in ICASSP , 2022
2022
Closest in time.
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. R. Glass, “SSAST: Self-Supervised Audio Spectrogram Transformer,” in AAAI , 2022
2022
Closest in time.
A. Baade, P. Peng, and D. Harwath, “MAE-AST: Masked autoencoding audio spectrogram transformer,” in Interspeech , 2022
2022
Closest in time.
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, T. N. Sainath, and S. Watanabe, “Self-supervised speech representation learning: A review,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1179–1210, 2022
2022
Closest in time.
Z. Tang, J. Cho, Y. Nie, and M. Bansal, “TVLT: Textless Vision-Language Transformer,” in NeurIPS , 2022
2022
Closest in time.
Y. Wu*, K. Chen*, T. Zhang*, Y. Hui*, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP , 2023
2023
Closest in time.
M. Sahidullah, T. Kinnunen, and C. Hanilçi, “A comparison of features for synthetic speech detection,” in INTERSPEECH 2015, 16th Annual Conference of the International Speech Communication Association, Dresden, Germany, September 6-10, 2015 . ISCA, 2015, pp. 2087–2091
2091
Closest in time.