Fetching the paper…
Reading the bibliography…
We learn audio representations by solving a novel self-supervised learning task, which consists of predicting the phase of the short-time Fourier transform from its magnitude.
“Signal estimation from modified short-time Fourier transform,”
D. Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Estimating and interpreting the instantaneous frequency of a signal. i. fundamentals,”
Boualem Boashash, · 1992
Earlier work this paper cites.
“Unsupervised feature learning for audio classification using convolutional deep belief networks,”
Honglak Lee, Yan Largman, Peter Pham, and Andrew Y Ng, · 2009
Earlier work this paper cites.
“Efficient Estimation of Word Representations in Vector Space,”
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean, · 2013
Earlier work this paper cites.
“Unsupervised Visual Representation Learning by Context Prediction,”
Carl Doersch, Abhinav Gupta, and Alexei A. Efros, · 2015
Earlier work this paper cites.
“LibriSpeech: An ASR Corpus Based on Public Domain Adio Books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus,”
David Snyder, Guoguo Chen, and Daniel Povey, · 2015
Earlier work this paper cites.
“Colorful Image Colorization,”
Richard Zhang, Phillip Isola, and Alexei A. Efros, · 2016
Earlier work this paper cites.
“Audio Word2Vec: Unsupervised Learning of Audio Segment Representations using Sequence-to-sequence Autoencoder,”
Yu-An Chung, Chao-Chung Wu, Chia-Hao Shen, Hung-Yi Lee, and Lin-Shan Lee, · 2016
Earlier work this paper cites.
“Context Encoders: Feature Learning by Inpainting,”
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros, · 2016
Earlier work this paper cites.
“CNN architectures for large-scale audio classification,”
Shawn Hershey, Sourish Chaudhuri, Daniel P.W. Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, Malcolm Slaney, Ron J Weiss, and Kevin Wilson, · 2017
Earlier work this paper cites.
“Audio Set: An ontology and human-labeled dataset for audio events,”
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter, · 2017
Cited alongside, same era.
“Unsupervised Feature Learning Based on Deep Models for Environmental Audio Tagging,”
Yong Xu, Qiang Huang, Wenwu Wang, Peter Foster, Siddharth Sigtia, Philip J B Jackson, and Mark D Plumbley, · 2017
Cited alongside, same era.
“Unsupervised Feature Learning for Audio Analysis,”
Matthias Meyer, Jan Beutel, and Lothar Thiele, · 2017
Cited alongside, same era.
“Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,”
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi, · 2017
Cited alongside, same era.
“Unsupervised Learning of Semantic Audio Representations,”
Aren Jansen, Manoj Plakal, Ratheet Pandya, Daniel P. W. Ellis, Shawn Hershey, Jiayang Liu, R. Channing Moore, and Rif A. Saurous, · 2018
Cited alongside, same era.
“Spoken Language Identification,” 2018
Tomasz Oponowicz, · 2018
Later among the works it cites.
“Automatic acoustic detection of birds through deep learning: the first Bird Audio Detection challenge,”
Dan Stowell, — Mike Wood, — Hanna Pamuła, Yannis Stylianou, and Hervé Glotin, · 2018
Later among the works it cites.
“Detection and Classification of Acoustic Scenes and Events,”
Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen, · 2018
Later among the works it cites.
“Self-supervised audio representation learning for mobile devices,”
Marco Tagliasacchi, Beat Gfeller, Félix de Chaumont Quitry, and Dominik Roblek, · 2019
Closest in time.
“Learning Problem-agnostic Speech Representations from Multiple Self-supervised Tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio, · 2019
Closest in time.
“Melnet: A generative model for audio in the frequency domain,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Unsupervised Representation Learning by Predicting Image Rotations,”
Spyros Gidaris, Praveer Singh, and Nikos Komodakis, · 2018
Cited alongside, same era.
“Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech,”
Yu-An Chung and James Glass, · 2018
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, and et al., · 2018
Cited alongside, same era.
“The conversation: Deep audio-visual speech enhancement,”
T. Afouras, J. S. Chung, and A. Zisserman, · 2018
Cited alongside, same era.
“Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition,”
Pete Warden, · 2018
Cited alongside, same era.
Sean Vasquez and Mike Lewis, · 2019
Closest in time.
“Adversarial Audio Synthesis,”
Chris Donahue, Julian McAuley, and Miller Puckette, · 2019
Closest in time.
“Waveglow: A flow-based generative network for speech synthesis,”
Ryan Prenger, Rafael Valle, and Bryan Catanzaro, · 2019
Closest in time.
“GANSynth: Adversarial Neural Audio Synthesis,”
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts, · 2019
Closest in time.
“Deep Griffin-Lim Iteration,”
Yoshiki Masuyama, Kohei Yatabe, Yuma Koizumi, Yasuhiro Oikawa, and Noboru Harada, · 2019
Closest in time.