Fetching the paper…
Reading the bibliography…
Currently, most speech processing techniques use magnitude spectrograms as front-end and are therefore by default discarding part of the signal: the phase.
Noise reduction using connectionist models
Shin’ichi Tamura and Alex Waibel · 1988
Earlier work this paper cites.
Speech enhancement based on a priori signal to noise estimation
Pascal Scalart et al · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Speech enhancement with missing data techniques using recurrent neural networks
Shahla Parveen and Phil Green · 2004
Earlier work this paper cites.
Evaluation of objective measures for speech enhancement
Yi Hu and Philipos C Loizou · 2006
Earlier work this paper cites.
Speech quality assessment
Philipos C Loizou · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Nonnegative signal factorization with learnt instrument models for sound source separation in close-microphone recordings
Julio J Carabias-Orti, Máximo Cobos, Pedro Vera-Candeas, and Francisco J Rodríguez-Serrano · 2013
Earlier work this paper cites.
Speech enhancement based on deep denoising autoencoder
Xugang Lu, Yu Tsao, Shigeki Matsuda, and Chiori Hori · 2013
Earlier work this paper cites.
The diverse environments multi-channel acoustic noise database: A database of multichannel environmental noise recordings
Joachim Thiemann, Nobutaka Ito, and Emmanuel Vincent · 2013
Earlier work this paper cites.
The voice bank corpus: Design, collection and data analysis of a large regional accent speech database
Christophe Veaux, Junichi Yamagishi, and Simon King · 2013
Earlier work this paper cites.
End-to-end learning for music audio
Sander Dieleman and Benjamin Schrauwen · 2014
Earlier work this paper cites.
Joint optimization of masks and deep recurrent neural networks for monaural source separation
Po-Sen Huang, Minje Kim, Mark Hasegawa-Johnson, and Paris Smaragdis · 2015
Earlier work this paper cites.
Convolutional neural networks-based continuous speech recognition using raw speech signal
Dimitri Palaz, Mathew Magimai Doss, and Ronan Collobert · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr
Felix Weninger, Hakan Erdogan, Shinji Watanabe, Emmanuel Vincent, Jonathan Le Roux, John R Hershey, and Björn Schuller · 2015
Cited alongside, same era.
A regression approach to speech enhancement based on deep neural networks
Yong Xu, Jun Du, Li-Rong Dai, and Chin-Hui Lee · 2015
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun · 2015
Cited alongside, same era.
Conditional image generation with pixelcnn decoders
Aäron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Later among the works it cites.
The microsoft 2016 conversational speech recognition system
W Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig · 2016
Later among the works it cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaolei Huang, Xiaogang Wang, and Dimitris Metaxas · 2016
Later among the works it cites.
Learning multiscale features directly from waveforms
Zhenyao Zhu, Jesse H Engel, and Awni Hannun · 2016
Later among the works it cites.
Monoaural audio source separation using deep convolutional neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Rishita Anubhai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, et al · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Speech enhancement in multiple-noise conditions using deep neural networks
Anurag Kumar and Dinei Florencio · 2016
Cited alongside, same era.
Samplernn: An unconditional end-to-end neural audio generation model
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Pixel recurrent neural networks
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Remixing music using source separation algorithms to improve the musical experience of cochlear implant users
Jordi Pons, Jordi Janer, Thilo Rode, and Waldo Nogueira · 2016
Cited alongside, same era.
Pritish Chandna, Marius Miron, Jordi Janer, and Emilia Gómez · 2017
Closest in time.
Deep cross-modal audio-visual generation
Lele Chen, Sudhanshu Srivastava, Zhiyao Duan, and Chenliang Xu · 2017
Closest in time.
Neural audio synthesis of musical notes with wavenet autoencoders
Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi · 2017
Closest in time.
Jongpil Lee and Juhan Nam · 2017
Closest in time.
Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam · 2017
Closest in time.
Segan: Speech enhancement generative adversarial network
Santiago Pascual, Antonio Bonafonte, and Joan Serra · 2017
Closest in time.
Speech enhancement using bayesian wavenet
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Florencio Dinei, and Mark Hasegawa-Johnson · 2017
Closest in time.
End-to-end source separation with adaptive front-ends
Shrikant Venkataramani and Paris Smaragdis · 2017
Closest in time.