Fetching the paper…
Reading the bibliography…
Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party.
“Multitalker Speech Separation With Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks”
Morten Kolbaek, Dong Yu, Zheng Tan and Jesper Jensen · 1913
Earlier work this paper cites.
“Speech analysis and synthesis by linear prediction of the speech wave”
Bishnu Atal and Suzanne Hanauer · 1971
Earlier work this paper cites.
“Auditory scene analysis: the perceptual organization of sound”, 1990
AS Bregman · 1990
Earlier work this paper cites.
“Auditory scene analysis: the perceptual organization of sound”, 1994, pp. 79–80
Albert Bregman · 1994
Earlier work this paper cites.
“A 2.4 kbit/s MELP coder candidate for the new US Federal Standard”
Alan McCree, Kwan Truong, E George, Thomas Barnwell and Vishu Viswanathan · 1996
Earlier work this paper cites.
“Single-channel speech separation using sparse non-negative matrix factorization”
Mikkel Schmidt and Rasmus Olsson · 2006
Earlier work this paper cites.
“Coupled dictionary training for exemplar-based speech enhancement”
Deepak Baby, Tuomas Virtanen and Tom Barker · 2014
Earlier work this paper cites.
“Exemplar-based speech enhancement for deep neural network based automatic speech recognition”
Deepak Baby, Jort Gemmeke and Tuomas Virtanen · 2015
Earlier work this paper cites.
“Coupled dictionaries for exemplar-based speech enhancement and automatic speech recognition”
Deepak Baby, Tuomas Virtanen and Jort Gemmeke · 2015
Earlier work this paper cites.
“Deep clustering: discriminative embeddings for segmentation and separation”
John Hershey, Zhuo Chen, Jonathan Le and Shinji Watanabe · 2016
Earlier work this paper cites.
“Single-channel multi-speaker separation using deep clustering”
Yusuf Isik, Jonathan Roux, Zhuo Chen, Shinji Watanabe and John. Hershey · 2016
Earlier work this paper cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation”
Dong Yu, Morten Kolbæk, Zheng-Hua Tan and Jesper Jensen · 2017
Earlier work this paper cites.
“Neural Discrete Representation Learning”
Aaron van Oord, Oriol Vinyals and Koray Kavukcuoglu · 2017
Cited alongside, same era.
“Noisy speech database for training speech enhancement algorithms and TTS models”, 2017
Cassia Valentini-Botinhao · 2017
Cited alongside, same era.
“TaSNet: time-domain audio separation network for real-time, single-channel speech separation”
Yi Luo and Nima Mesgarani · 2018
Cited alongside, same era.
“Wavenet based low rate speech coding”
W Kleijn et al · 2018
Cited alongside, same era.
“Spectral Feature Mapping with Mimic Loss for Robust Speech Recognition”
Deblin Bagchi, Peter Plantinga, Adam Stiff and Eric Fosler-Lussier · 2018
Cited alongside, same era.
“Speech Processing for Digital Home Assistants: Combining Signal Processing With Deep-Learning Techniques”
“Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation”
Yi Luo, Zhuo Chen and Takuya Yoshioka · 2020
Later among the works it cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis”
Jungil Kong, Jaehyeon Kim and Jaekyoung Bae · 2020
Later among the works it cites.
“Robust low rate speech coding based on cloned networks and wavenet”
Felicia Lim, W Kleijn, Michael Chinen and Jan Skoglund · 2020
Later among the works it cites.
“Discretalk: Text-to-speech as a machine translation problem”
Tomoki Hayashi and Shinji Watanabe · 2020
Later among the works it cites.
“Libri-light: A benchmark for asr with limited or no supervision”
Jacob Kahn et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinhold Haeb-Umbach et al · 2019
Cited alongside, same era.
“Speech enhancement using end-to-end speech recognition objectives”
Aswin Subramanian et al · 2019
Cited alongside, same era.
“MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis”
Kundan Kumar et al · 2019
Cited alongside, same era.
“A comparative study on transformer vs rnn in speech applications”
Shigeki Karita et al · 2019
Cited alongside, same era.
“MOSNet: Deep Learning based Objective Assessment for Voice Conversion”
Chen-Chou Lo et al · 2019
Cited alongside, same era.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation”
Yi Luo and Nima Mesgarani · 2019
Cited alongside, same era.
Hervé Bredin et al · 2020
Later among the works it cites.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units”
Wei-Ning Hsu et al · 2021
Closest in time.
“SUPERB: Speech processing Universal PERformance Benchmark”
Shu-wen Yang et al · 2021
Closest in time.
“Speech Resynthesis from Discrete Disentangled Self-Supervised Representations”
Adam Polyak et al · 2021
Closest in time.
“Direct speech-to-speech translation with discrete units”
Ann Lee et al · 2021
Closest in time.
“MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement”
Szu-Wei Fu et al · 2021
Closest in time.