Fetching the paper…
Reading the bibliography…
We explore the possibility of leveraging accelerometer data to perform speech enhancement in very noisy conditions.
1907
Earlier work this paper cites.
M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” in
1979
Earlier work this paper cites.
P. Scalart and J. V. Filho, “Speech enhancement based on a priori signal to noise estimation,” in
1996
Earlier work this paper cites.
H. Attias, J. C. Platt, A. Acero, and L. Deng, “Speech denoising and dereverberation using probabilistic models,” in
2001
Earlier work this paper cites.
S. Kamath and P. Loizou, “A multi-band spectral subtraction method for enhancing speech corrupted by colored noise,” in
2002
Earlier work this paper cites.
N. Zeghidour and D. Grangier, “Wavesplit: End-to-end speech separation by speaker clustering,”
2002
Earlier work this paper cites.
J. R. Hershey, T. T. Kristjansson, and Z. Zhang, “Model-based fusion of bone and air sensors for speech enhancement and robust speech recognition,” in
2004
Earlier work this paper cites.
P. C. Loizou, “Speech enhancement based on perceptually motivated bayesian estimators of the magnitude spectrum,”
2005
Earlier work this paper cites.
A. M. Reddy and B. Raj, “Soft mask methods for single-channel speaker separation,”
2007
Earlier work this paper cites.
E. M. Grais and H. Erdogan, “Single channel speech music separation using nonnegative matrix factorization and spectral masks,” in
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in
2013
Cited alongside, same era.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in
2013
Cited alongside, same era.
X. Feng, Y. Zhang, and J. Glass, “Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition,” in
2014
Cited alongside, same era.
Y. Michalevsky, D. Boneh, and G. Nakibly, “Gyrophone: recognizing speech from gyroscope signals,” in
2014
Cited alongside, same era.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr,” in
2015
Cited alongside, same era.
D. Rethage, J. Pons, and X. Serra, “A wavenet for speech denoising,” in
2018
Later among the works it cites.
Y. Luo and N. Mesgarani, “TaSNet: Time-domain audio separation network for real-time, single-channel speech separation,” in
2018
Later among the works it cites.
J. Hou, S. Wang, Y. Lai, Y. Tsao, H. Chang, and H. Wang, “Audio-visual speech enhancement using multimodal deep convolutional neural networks,”
2018
Later among the works it cites.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein, “Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,”
2018
Later among the works it cites.
H. Liu, Y. Tsao, and C. Fuh, “Bone-conducted speech enhancement using deep denoising autoencoder,”
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Cited alongside, same era.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in
2016
Cited alongside, same era.
S. H. Djork-Arné Clevert, Thomas Unterthiner, “Fast and accurate deep network learning by exponential linear units (ELUs),” in
2016
Cited alongside, same era.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” in
2017
Cited alongside, same era.
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
2017
Cited alongside, same era.
C. Donahue, B. Li, and R. Prabhavalkar, “Exploring speech enhancement with generative adversarial networks for robust speech recognition,” in
2018
Cited alongside, same era.
S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous, “Differentiable consistency constraints for improved deep speech enhancement,” in
2019
Later among the works it cites.
J. L. Roux, G. Wichern, S. Watanabe, A. M. Sarroff, and J. R. Hershey, “The Phasebook: Building complex masks via discrete representations for source separation,” in
2019
Later among the works it cites.
I. Kavalerov, S. Wisdom, H. Erdogan, B. Patton, K. W. Wilson, J. L. Roux, and J. R. Hershey, “Universal sound separation,” in
2019
Later among the works it cites.
T. Ochiai, M. Delcroix, K. Kinoshita, A. Ogawa, and T. Nakatani, “Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues,” in
2019
Later among the works it cites.
I. Ariav and I. Cohen, “An end-to-end multimodal voice activity detection using wavenet encoder and residual networks,”
2019
Later among the works it cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in
2019
Later among the works it cites.