Fetching the paper…
Reading the bibliography…
In this paper, we present a framework for contrastive learning for audio representations, in a self supervised frame work without access to any ground truth labels.
“Size matters: An empirical study of neural network training for large vocabulary continuous speech recognition,”
Dan Ellis and Nelson Morgan, · 1999
Earlier work this paper cites.
“Reducing the dimensionality of data with neural networks,”
Geoffrey E Hinton and Ruslan R Salakhutdinov, · 2006
Earlier work this paper cites.
“Efficient estimation of word representations in vector space,”
Tomas Mikolov, Kai Chen, Gregory S. Corrado, and Jeffrey Dean, · 2013
Earlier work this paper cites.
“A dataset and taxonomy for urban sound research,”
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello, · 2014
Earlier work this paper cites.
“Unsupervised representation learning with deep convolutional generative adversarial networks,”
Alec Radford, Luke Metz, and Soumith Chintala, · 2015
Earlier work this paper cites.
“librosa: Audio and music signal analysis in python,”
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto, · 2015
Earlier work this paper cites.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Earlier work this paper cites.
“Cnn architectures for large-scale audio classification,”
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, · 2017
Cited alongside, same era.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Cited alongside, same era.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Neural style transfer for audio spectograms,”
Prateek Verma and Julius O Smith, · 2018
Cited alongside, same era.
“Neuralogram: A deep neural network based representation for audio signals,”
Prateek Verma, Chris Chafe, and Jonathan Berger, · 2019
Later among the works it cites.
“Speechbert: Cross-modal pre-trained language model for end-to-end spoken question answering,”
Yung-Sung Chuang, Chi-Liang Liu, and Hung-Yi Lee, · 2019
Later among the works it cites.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Closest in time.
Accessed: 2020–10-19
“Sentiment neuron open ai,” https://openai.com/blog/unsupervised-sentiment-neuron/ , · 2020
Closest in time.
“Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision,”
Aren Jansen, Daniel PW Ellis, Shawn Hershey, R Channing Moore, Manoj Plakal, Ashok C Popat, and Rif A Saurous, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Audio-linguistic embeddings for spoken sentences,”
Albert Haque, Michelle Guo, Prateek Verma, and Li Fei-Fei, · 2019
Cited alongside, same era.
“A software framework for musical data augmentation.,”
Brian McFee, Eric J Humphrey, and Juan Pablo Bello,
Cited in the paper.
“Exploring data augmentation for improved singing voice detection with neural networks.,”
Jan Schlüter and Thomas Grill,
Cited in the paper.
Closest in time.
“A deep learning approach for low-latency packet loss concealment of audio signals in networked music performance applications,”
Prateek Verma, Alessandro Ilic Mezzay, Chris Chafe, and Cristina Rottondi, · 2020
Closest in time.
“Multi-task self-supervised visual learning,”
Carl Doersch and Andrew Zisserman, · 2060
Closest in time.