Fetching the paper…
Reading the bibliography…
Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today.
Multiclass language identification using deep learning on spectral images of audio signals
Shauna Revay and Matthew Teschke · 1905
Earlier work this paper cites.
The relation of pitch to frequency: A revised scale
Stanley S Stevens and John Volkmann · 1940
Earlier work this paper cites.
Theory of communication. part 1: The analysis of information
Dennis Gabor · 1946
Earlier work this paper cites.
Elements of psychophysics , volume 1
Gustav Theodor Fechner, Davis H Howes, and Edwin Garrigues Boring · 1966
Earlier work this paper cites.
Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
Steven Davis and Paul Mermelstein · 1980
Earlier work this paper cites.
Speech communication : human and machine
D. O’Shaughnessy · 1987
Earlier work this paper cites.
An Introduction to the Bootstrap
Bradley Efron and Robert J. Tibshirani · 1993
Earlier work this paper cites.
The mel scale’s disqualifying bias and a consistency of pitch-difference equisections in 1956 with equal cochlear distances and equal frequency ratios
Donald D. Greenwood · 1997
Earlier work this paper cites.
Fitting the mel scale
S. Umesh, L. Cohen, and D. Nelson · 1999
Earlier work this paper cites.
Automatic speech recognition: An auditory perspective
Nelson Mogran, Hervé Bourlard, and Hynek Hermansky · 2004
Earlier work this paper cites.
Gammatone features and feature combination for large vocabulary speech recognition
Ralf Schluter, Ilja Bezrukov, Hermann Wagner, and Hermann Ney · 2007
Earlier work this paper cites.
Effect of compressing the dynamic range of the power spectrum in modulation filtering based speech enhancement
James G Lyons and Kuldip K Paliwal · 2008
Earlier work this paper cites.
Hilbert envelope based features for far-field speech recognition
Samuel Thomas, Sriram Ganapathy, and Hynek Hermansky · 2008
Earlier work this paper cites.
Modulation spectral features for robust far-field speaker identification
Tiago H Falk and Wai-Yip Chan · 2009
Earlier work this paper cites.
Learning a better representation of speech soundwaves using restricted boltzmann machines
Navdeep Jaitly and Geoffrey Hinton · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Dimitri Palaz, Ronan Collobert, and Mathew Magimai Doss · 2013
Earlier work this paper cites.
Learning filter banks within a deep neural network framework
T. Sainath, Brian Kingsbury, Abdel rahman Mohamed, and B. Ramabhadran · 2013
Earlier work this paper cites.
Deep scattering spectrum
Joakim Andén and Stéphane Mallat · 2014
Earlier work this paper cites.
CREMA-D: Crowd-sourced emotional multimodal actors dataset
Houwei Cao, David G Cooper, Michael K Keutmann, Ruben C Gur, Ani Nenkova, and Ragini Verma · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Speech acoustic modeling from raw multichannel waveforms
Yedid Hoshen, Ron Weiss, and Kevin W Wilson · 2015
Cited alongside, same era.
Convolutional neural networks-based continuous speech recognition using raw speech signal
Dimitri Palaz, Mathew Magimai Doss, and Ronan Collobert · 2015
Cited alongside, same era.
Learning the speech front-end with raw waveform cldnns
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals · 2015
Cited alongside, same era.
Acoustic modelling with cd-ctc-smbr lstm rnns
TUT Urban Acoustic Scenes 2018, Development dataset, April 2018
Toni Heittola, Annamaria Mesaros, and Tuomas Virtanen · 2018
Later among the works it cites.
Birdvox-full-night: A dataset and benchmark for avian flight call detection
Vincent Lostanlen, Justin Salamon, Andrew Farnsworth, Steve Kelling, and Juan Pablo Bello · 2018
Later among the works it cites.
Speaker recognition from raw waveform with sincnet
Mirco Ravanelli and Yoshua Bengio · 2018
Later among the works it cites.
Zero-mean convolutions for level-invariant singing voice detection
Jan Schlüter and Bernhard Lehner · 2018
Later among the works it cites.
Dan Stowell, — Mike Wood, — Hanna Pamuła, Yannis Stylianou, and Hervé Glotin · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew Senior, Hasim Sak, Felix de Chaumont Quitry, Tara N. Sainath, and Kanishka Rao · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Now Playing: Continuous low-power music recognition
Blaise Agüera y Arcas, Beat Gfeller, Ruiqi Guo, Kevin Kilgour, Sanjiv Kumar, James Lyon, Julian Odell, Marvin Ritter, Dominik Roblek, Matthew Sharifi, and Mihajlo Velimirović · 2017
Cited alongside, same era.
Reducing bias in production speech models
Eric Battenberg, Rewon Child, Adam Coates, Christopher Fougner, Yashesh Gaur, Jiaji Huang, Heewoo Jun, Ajay Kannan, Markus Kliegl, Atul Kumar, et al · 2017
Cited alongside, same era.
Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Cited alongside, same era.
Pete Warden · 2018
Later among the works it cites.
Learning filterbanks from raw speech for phone recognition
Neil Zeghidour, Nicolas Usunier, Iasonas Kokkinos, Thomas Schatz, Gabriel Synnaeve, and Emmanuel Dupoux · 2018
Later among the works it cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, T. Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley · 2019
Later among the works it cites.
Per-channel energy normalization: Why and how
V. Lostanlen, J. Salamon, M. Cartwright, B. McFee, A. Farnsworth, S. Kelling, and J. P. Bello · 2019
Later among the works it cites.
Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani · 2019
Later among the works it cites.
Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation
Yi Luo, Zhuo Chen, and Takuya Yoshioka · 2019
Later among the works it cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Later among the works it cites.
End-to-end asr: from supervised to semi-supervised learning with modern architectures
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Edouard Grave, Tatiana Likhomanenko, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Later among the works it cites.
Making convolutional networks shift-invariant again
Richard Zhang · 2019
Later among the works it cites.
Cgcnn: Complex gabor convolutional neural network on raw speech
Paul-Gauthier Noé, Titouan Parcollet, and Mohamed Morchid · 2020
Later among the works it cites.
Filterbank design for end-to-end speech separation
Manuel Pariente, Samuele Cornell, Antoine Deleforge, and Emmanuel Vincent · 2020
Later among the works it cites.
State-of-the-art speaker recognition with neural network embeddings in nist sre18 and speakers in the wild evaluations
Jesús Villalba, Nanxin Chen, David Snyder, Daniel Garcia-Romero, Alan McCree, Gregory Sell, Jonas Borgstrom, Leibny Paola García-Perera, Fred Richardson, Réda Dehak, et al · 2020
Later among the works it cites.