Fetching the paper…
Reading the bibliography…
We present a multimodal framework to learn general audio representations from videos.
J. Bromley, J. W. Bentz, L. Bottou, I. Guyon, Y. LeCun, C. Moore, E. Säckinger, and R. Shah, “Signature verification using a “siamese” time delay neural network,” International Journal of Pattern Recognition and Artificial Intelligence , vol. 7, no. 04, pp. 669–688, 1993
1993
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 5206–5210, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” IEEE International Conference on Computer Vision (ICCV) , pp. 609–617, 2017
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “CNN architectures for large-scale audio classification,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 131–135, 2017
2017
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio Set: An ontology and human-labeled dataset for audio events,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 776–780, 2017
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with WaveNet autoencoders,” International Conference on Machine Learning (ICML) , vol. 70, pp. 1068–1077, 06–11 Aug 2017. [Online]. Available: http://proceedings.mlr.press/v70/engel17a.html
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Jansen, M. Plakal, R. Pandya, D. P. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, “Unsupervised learning of semantic audio representations,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 126–130, 2018
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
T. Heittola, A. Mesaros, and T. Virtanen, “TUT urban acoustic scenes 2018, development dataset,” 2018
2018
Cited alongside, same era.
K. MacLean, “VoxForge,” Ken MacLean.[Online]. Available: http://www.voxforge.org/home.[Acedido em 2012] , 2018
2018
Cited alongside, same era.
P. Bachman, R. D. Hjelm, and W. Buchwalter, “Learning representations by maximizing mutual information across views,” Advances in Neural Information Processing Systems (NeurIPS) , pp. 15 509–15 519, 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , pp. 4171–4186, 2019
J.-B. Alayrac, A. Recasens, R. Schneider, R. Arandjelović, J. Ramapuram, J. De Fauw, L. Smaira, S. Dieleman, and A. Zisserman, “Self-supervised multimodal versatile networks,” Advances in Neural Information Processing Systems (NeurIPS) , pp. 25–37, 2020
2020
Later among the works it cites.
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord, “Learning robust and multilingual speech representations,” Conference on Empirical Methods in Natural Language Processing (EMNLP) , pp. 1182–1192, 2020
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in Neural Information Processing Systems (NeurIPS) , pp. 12 449–12 460, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” Interspeech , pp. 3465–3469, 2019
2019
Cited alongside, same era.
D. Stowell, M. D. Wood, H. Pamuła, Y. Stylianou, and H. Glotin, “Automatic acoustic detection of birds through deep learning: the first bird audio detection challenge,” Methods in Ecology and Evolution , vol. 10, no. 3, pp. 368–380, 2019
2019
Cited alongside, same era.
J. Lin, C. Gan, and S. Han, “TSM: Temporal shift module for efficient video understanding,” IEEE/CVF International Conference on Computer Vision (ICCV) , pp. 7083–7093, 2019
2019
Cited alongside, same era.
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” International Conference on Machine Learning (ICML) , pp. 6105–6114, 2019
2019
Cited alongside, same era.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” International Conference on Machine Learning (ICML) , pp. 1597–1607, 2020
2020
Cited alongside, same era.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 9729–9738, 2020
2020
Cited alongside, same era.
O. J. Hénaff, A. Srinivas, J. De Fauw, A. Razavi, C. Doersch, S. Eslami, and A. van den Oord, “Data-efficient image recognition with contrastive predictive coding,” International Conference on Machine Learning (ICML) , pp. 4182–4192, 2020
2020
Cited alongside, same era.
2020
Later among the works it cites.
L. Wang and A. van den Oord, “Multi-format contrastive learning of audio representations,” NeurIPS Workshops (Self-Supervised Learning for Speech and Audio Processing) , 2020
2020
Later among the works it cites.
A. Jansen, D. P. Ellis, S. Hershey, R. C. Moore, M. Plakal, A. C. Popat, and R. A. Saurous, “Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 121–125, 2020
2020
Later among the works it cites.
M. Tagliasacchi, B. Gfeller, F. de Chaumont Quitry, and D. Roblek, “Pre-training audio representations with self-supervision,” IEEE Signal Processing Letters , vol. 27, pp. 600–604, 2020
2020
Later among the works it cites.
L. Wang, K. Kawakami, and A. van den Oord, “Contrastive predictive coding of audio with an adversary,” Interspeech , pp. 826–830, 2020
2020
Later among the works it cites.
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de Chaumont Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, “Towards learning a universal non-semantic representation of speech,” Interspeech , pp. 140–144, 2020
2020
Later among the works it cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNS: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2020
2020
Later among the works it cites.
2021
Closest in time.
N. Zeghidour, O. Teboul, F. d. C. Quitry, and M. Tagliasacchi, “LEAF: A learnable frontend for audio classification,” International Conference on Learning Representations (ICLR) , 2021
2021
Closest in time.