Fetching the paper…
Reading the bibliography…
Neural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence.
Some experiments on the recognition of speech, with one and with two ears
E. Colin Cherry · 1953
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser · 1989
Earlier work this paper cites.
Auditory scene analysis: the perceptual organization of sound
Albert S Bregman · 1994
Earlier work this paper cites.
Blind source separation-semiparametric statistical approach
Shun-Ichi Amari and J.-F. Cardoso · 1997
Earlier work this paper cites.
Iterative and sequential kalman filter-based speech enhancement algorithms
Sharon Gannot, David Burshtein, and Ehud Weinstein · 1998
Earlier work this paper cites.
Extended ICA removes artifacts from electroencephalographic recordings
Tzyy-Ping Jung, Colin Humphries, Te-Won Lee, Scott Makeig, Martin J McKeown, Vicente Iragui, and Terrence J Sejnowski · 1998
Earlier work this paper cites.
Informational and energetic masking effects in the perception of two simultaneous talkers
Douglas S Brungart · 2001
Earlier work this paper cites.
One microphone source separation
Sam T Roweis · 2001
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
An overview of automatic speaker diarization systems
Sue E Tranter and Douglas A Reynolds · 2006
Earlier work this paper cites.
Speaker diarization: a review of recent research
Xavier Anguera, Simon Bozonnet, Nicholas Evans, Corinne Fredouille, Gerald Friedland, and Oriol Vinyals · 2012
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Deep neural networks for single-channel multi-talker speech recognition
Chao Weng, Dong Yu, Michael L Seltzer, and Jasha Droppo · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville · 2015
Cited alongside, same era.
Listen, attend and spell: a neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Cited alongside, same era.
Deep clustering: discriminative embeddings for segmentation and separation
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe · 2016
Cited alongside, same era.
Single-channel multi-speaker separation using deep clustering
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R. Hershey · 2016
Cited alongside, same era.
Google’s neural machine translation system: bridging the gap between human and machine translation
End-to-end multi-speaker speech recognition
Shane Settle, Jonathan Le Roux, Takaaki Hori, Shinji Watanabe, and John R Hershey · 2018
Later among the works it cites.
Listen, think and listen again: capturing top-down auditory attention for speaker-independent speech separation
Jing Shi, Jiaming Xu, Guangcan Liu, and Bo Xu · 2018
Later among the works it cites.
Espnet: end-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al · 2018
Later among the works it cites.
End-to-end monaural multi-speaker asr system without pretraining
Xuankai Chang, Yanmin Qian, Kai Yu, and Shinji Watanabe · 2019
Later among the works it cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
Informational masking in speech recognition
Gerald Kidd and H Steven Colburn · 2017
Cited alongside, same era.
Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks
Morten Kolbaek, Dong Yu, Zheng Hua Tan, Jesper Jensen, Morten Kolbaek, Dong Yu, Zheng Hua Tan, and Jesper Jensen · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
End-to-end speech recognition with word-based rnn language models
Takaaki Hori, Jaejin Cho, and Shinji Watanabe · 2018
Cited alongside, same era.
Listening to each speaker one by one with recurrent selective hearing networks
Keisuke Kinoshita, Lukas Drude, Marc Delcroix, and Tomohiro Nakatani · 2018
Cited alongside, same era.
Speaker-independent speech separation with deep attractor network
Yi Luo, Zhuo Chen, and Nima Mesgarani · 2018
Cited alongside, same era.
Tobias Menne, Ilya Sklyar, Ralf Schlüter, and Hermann Ney · 2019
Later among the works it cites.
Speech enhancement using end-to-end speech recognition objectives
Aswin Shanmugam Subramanian, Xiaofei Wang, Murali Karthick Baskar, Shinji Watanabe, Toru Taniguchi, Dung Tran, and Yuya Fujita · 2019
Later among the works it cites.
Recursive Speech Separation for Unknown Number of Speakers
Naoya Takahashi, Sudarsanam Parthasaarathy, Nabarun Goswami, and Yuki Mitsufuji · 2019
Later among the works it cites.
Neural speaker diarization with speaker-wise chain rule
Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Yawen Xue, Jing Shi, and Kenji Nagamatsu · 2020
Closest in time.
Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation
Yi Luo, Zhuo Chen, and Takuya Yoshioka · 2020
Closest in time.
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Khokhlov, Mariya Korenevskaya, Ivan Sorokin, Tatiana Timofeeva, Anton Mitrofanov, Andrei Andrusenko, Ivan Podluzhny, et al · 2020
Closest in time.
Voice separation with an unknown number of multiple speakers
Eliya Nachmani, Yossi Adi, and Lior Wolf · 2020
Closest in time.
Chandan KA Reddy, Ebrahim Beyrami, Harishchandra Dubey, Vishak Gopal, Roger Cheng, Ross Cutler, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, et al · 2020
Closest in time.
Far-field location guided target speech extraction using end-to-end speech recognition objectives
Aswin Shanmugam Subramanian, Chao Weng, Meng Yu, Shi-Xiong Zhang, Yong Xu, Shinji Watanabe, and Dong Yu · 2020
Closest in time.
Weighted speech distortion losses for neural-network-based real-time speech enhancement
Yangyang Xia, Sebastian Braun, Chandan K. A. Reddy, Harishchandra Dubey, Ross Cutler, and Ivan Tashev · 2020
Closest in time.