Fetching the paper…
Reading the bibliography…
Inspired by the recent progress in self-supervised learning for computer vision that generates supervision using data augmentations, we explore a new general-purpose audio representation learning approach.
2006
Earlier work this paper cites.
U. Zölzer, DAFX: Digital Audio Effects . John Wiley & Sons, 2011
2011
Earlier work this paper cites.
G. C. Calafiore and L. El Ghaoui, Optimization Models . Cambridge University Press, 2014
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in ACM-MM’14 , Orlando, FL, USA, Nov. 2014, pp. 1041–1044
2014
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in ICML , 2015, pp. 448–456
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in ICASSP , 2017, pp. 776–780
2017
Earlier work this paper cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with WaveNet autoencoders,” in ICML , 2017, pp. 1068–1077
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Proc. Interspeech 2017 , 2017, pp. 2616–2620
2017
Earlier work this paper cites.
A. Jansen, M. Plakal, R. Pandya, D. P. W. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, “Unsupervised learning of semantic audio representations,” in ICASSP , 2018, pp. 126–130
2018
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” in ICLR , 2018. [Online]. Available: https://openreview.net/forum?id=r1Ddp1-Rb
2018
Earlier work this paper cites.
Z. Zhang, S. Xu, S. Cao, and S. Zhang, “Deep convolutional neural network with mixup for environmental sound classification,” in PRCV , 2018, pp. 356–367
2018
Earlier work this paper cites.
K. MacLean, “Voxforge,” 2018. [Online]. Available: http://www.voxforge.org/home
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Cited alongside, same era.
J. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, “Look, listen and learn more: Design choices for deep audio embeddings,” in ICASSP , Brighton, UK, May 2019, pp. 3852–â3 856
2019
Cited alongside, same era.
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in ICML , 2019, pp. 6105–6114
2019
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in ICML , 2020, pp. 1597–1607
2020
Later among the works it cites.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-task self-supervised learning for robust speech recognition,” in ICASSP , 2020, pp. 6989–6993
2020
Later among the works it cites.
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra, “Unsupervised Contrastive Learning of Sound Event Representations,” arXiv preprint arXiv::2011.07616 , 2020
2020
Later among the works it cites.
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de C. Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, “Towards learning a universal non-semantic representation of speech,” arXiv preprint arXiv::2002.12764 , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in SIGKDD , 2019
2019
Cited alongside, same era.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in NeurIPS , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
P. H. Le-Khac, G. Healy, and A. F. Smeaton, “Contrastive representation learning: A framework and review,” IEEE Access , vol. 8, pp. 193 907–193 934, 2020
2020
Cited alongside, same era.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020, pp. 9726–9735
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive learning of general-purpose audio representations,” arXiv preprint arXiv::2010.10915 , 2020
2020
Later among the works it cites.
V. Verma, M.-T. Luong, K. Kawaguchi, H. Pham, and Q. V. Le, “Towards domain-agnostic contrastive learning,” arXiv preprint arXiv::2011.04419 , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Y. Koizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “The NTT DCASE2020 challenge task 6 system: Automated audio captioning with keywords and sentence length estimation,” DCASE2020 Challenge, Tech. Rep., 2020
2020
Later among the works it cites.
D. Takeuchi, Y. Koizumi, Y. Ohishi, N. Harada, and K. Kashino, “Effects of word-frequency based pre- and post- processings for audio captioning,” in DCASE2020 , 2020, pp. 190–194
2020
Later among the works it cites.