Fetching the paper…
Reading the bibliography…
In this work, we investigate if the learned encoder of the end-to-end convolutional time domain audio separation network (Conv-TasNet) is the key to its recent success, or if the encoder can just as well be replaced by a deterministic hand-crafted filterbank.
“A Generalized Inverse for Matrices,”
R. Penrose, · 1955
Earlier work this paper cites.
“An Efficient Auditory Filterbank based on the Gammatone Function,”
R. Patterson, Ian Nimmo-Smith, J. Holdsworth, and P. Rice, · 1988
Earlier work this paper cites.
“Derivation of Auditory Filter Shapes from Notched-Noise Data,”
B. R. Glasberg and B. C. J. Moore, · 1990
Earlier work this paper cites.
“Complex Sounds and Auditory Images,”
R.D. Patterson, K. Robinson, J. Holdsworth, D. McKeown, C. Zhang, and M. Allerhand, · 1992
Earlier work this paper cites.
“Frequency Analysis and Synthesis using a Gammatone Filterbank,”
V. Hohmann, · 2002
Earlier work this paper cites.
“Deep clustering: Discriminative Embeddings for Segmentation and Separation,”
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, · 2016
Earlier work this paper cites.
“Permutation Invariant Training of Deep Models for Speaker-Independent Multi-Talker Speech Separation,”
D. Yu, M. Kolbaek, Z. Tan, and J. Jensen, · 2017
Earlier work this paper cites.
“Multitalker Speech Separation With Utterance-Level Permutation Invariant Training of Deep Recurrent Neural Networks,”
M. Kolbæk, D. Yu, Z. Tan, and J. Jensen, · 2017
Cited alongside, same era.
“Deep Attractor Network for Single-Microphone Speaker Separation,”
Z. Chen, Y. Luo, and N. Mesgarani, · 2017
Cited alongside, same era.
“Temporal Convolutional Networks for Action Segmentation and Detection,”
C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, · 2017
Cited alongside, same era.
“Alternative Objective Functions for Deep Clustering,”
Z. Wang, J. Le Roux, and J. R. Hershey, · 2018
Cited alongside, same era.
“TaSNet: Time-Domain Audio Separation Network for Real-Time, Single-Channel Speech Separation,”
Y. Luo and N. Mesgarani, · 2018
Cited alongside, same era.
“Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation,”
“End-to-End Monaural Speech Separation with Multi-Scale Dynamic Weighted Gated Dilated Convolutional Pyramid Network,”
Z. Shi, H. Lin, L. Liu, R. Liu, S. Hayakawa, S. Harada, and J. Han, · 2019
Closest in time.
“Deep Attention Gated Dilated Temporal Convolutional Networks with Intra-Parallel Convolutional Modules for End-to-End Monaural Speech Separation,”
Z. Shi, H. Lin, L. Liu, R. Liu, J. Han, and A. Shi, · 2019
Closest in time.
“Demystifying TasNet: A Dissecting Approach,”
J. Heitkaemper, D. Jakobeit, C. Boeddeker, L. Drude, and R. Haeb-Umbach, · 2019
Closest in time.
“End-To-End Training of Time Domain Audio Separation and Recognition,”
T. von Neumann, K. Kinoshita, L. Drude, C. Boeddeker, M. Delcroix, T. Nakatani, and R. Haeb-Umbach, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Luo and N. Mesgarani, · 2019
Cited alongside, same era.
M. Pariente, S. Cornell, A. Deleforge, and E. Vincent, · 2019
Closest in time.
“Influence of Speaker-Specific Parameters on Speech Separation Systems,”
D. Ditter and T. Gerkmann, · 2019
Closest in time.