Fetching the paper…
Reading the bibliography…
We propose the Neuralogram -- a deep neural network based representation for understanding audio signals which, as the name suggests, transforms an audio signal to a dense, compact representation based upon embeddings learned via a neural architecture.
“A duplex theory of pitch perception,”
Joseph Carl Robnett Licklider, · 1951
Earlier work this paper cites.
“Correlograms and the separation of sounds,”
Richard O Duda, Richard F Lyon, and Malcolm Slaney, · 1990
Earlier work this paper cites.
“Calculation of a constant q spectral transform,”
Judith C Brown, · 1991
Earlier work this paper cites.
“Rasta processing of speech,”
Hynek Hermansky and Nelson Morgan, · 1994
Earlier work this paper cites.
“To catch a chorus: Using chroma-based representations for audio thumbnailing,”
Mark A Bartsch and Gregory H Wakefield, · 2001
Earlier work this paper cites.
“The lower limit of melodic pitch,”
Daniel Pressnitzer, Roy D Patterson, and Katrin Krumbholz, · 2001
Earlier work this paper cites.
“Multiple scale music segmentation using rhythm, timbre, and harmony,”
Kristoffer Jensen, · 2007
Earlier work this paper cites.
“Music structure analysis using a probabilistic fitness measure and a greedy search algorithm,”
Jouni Paulus and Anssi Klapuri, · 2009
Earlier work this paper cites.
“Pitch perception,”
William A Yost, · 2009
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks,”
Xavier Glorot and Yoshua Bengio, · 2010
Earlier work this paper cites.
“Distributed representations of words and phrases and their compositionality,”
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean, · 2013
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Dropout: a simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Cited alongside, same era.
“Structural segmentation of hindustani concert audio with posterior features,”
Prateek Verma, TP Vinutha, Parthe Pandit, and Preeti Rao, · 2015
Cited alongside, same era.
“A neural algorithm of artistic style,”
Leon A Gatys, Alexander S Ecker, and Matthias Bethge, · 2015
Cited alongside, same era.
“Learning the speech front-end with raw waveform cldnns,”
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals, · 2015
Cited alongside, same era.
“Learning multiscale features directly fromwaveforms,”
Zhenyao Zhu, Jesse H Engel, and Awni Y Hannun, · 2016
Later among the works it cites.
“CNN architectures for large-scale audio classification,”
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al., · 2017
Later among the works it cites.
“Neural audio synthesis of musical notes with wavenet autoencoders,”
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi, · 2017
Later among the works it cites.
“A neural representation of sketch drawings,”
David Ha and Douglas Eck, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Wavenet: A generative model for raw audio,”
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Cited alongside, same era.
Allen Huang and Raymond Wu, · 2016
Cited alongside, same era.
“Frequency Estimation from Waveforms Using Multi-Layered Neural Networks.,”
Prateek Verma and Ronald W Schafer, · 2016
Cited alongside, same era.
“node2vec: Scalable feature learning for networks,”
Aditya Grover and Jure Leskovec, · 2016
Cited alongside, same era.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Cited alongside, same era.
Albert Haque, Michelle Guo, and Prateek Verma, · 2018
Later among the works it cites.
“Neural style transfer for audio spectograms,”
Prateek Verma and Julius O Smith, · 2018
Later among the works it cites.
“Deep contextualized word representations,”
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Later among the works it cites.
“Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech,”
Yu-An Chung and James Glass, · 2018
Later among the works it cites.
“Speech commands: A dataset for limited-vocabulary speech recognition,”
Pete Warden, · 2018
Later among the works it cites.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Later among the works it cites.
“Audio-linguistic embeddings for spoken sentences,”
Albert Haque, Michelle Guo, Prateek Verma, and Li Fei-Fei, · 2019
Closest in time.