Fetching the paper…
Reading the bibliography…
An important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones.
“Image method for efficiently simulating small-room acoustics,”
Jont B. Allen and David A. Berkley, · 1979
Earlier work this paper cites.
“Improved MVDR beamforming using single-channel mask prediction networks.,”
Hakan Erdogan, John R. Hershey, Shinji Watanabe, Michael I. Mandel, and Jonathan Le Roux, · 1985
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“A study of learning based beamforming methods for speech recognition,”
Xiong Xiao, Chenglin Xu, Zhaofeng Zhang, Shengkui Zhao, Sining Sun, Shinji Watanabe, Longbiao Wang, Lei Xie, Douglas L Jones, Eng Siong Chng, et al., · 2016
Earlier work this paper cites.
“Neural network based spectral mask estimation for acoustic beamforming,”
Jahn Heymann, Lukas Drude, and Reinhold Haeb-Umbach, · 2016
Earlier work this paper cites.
“Beamforming networks using spatial covariance features for far-field speech recognition,”
Xiong Xiao, Shinji Watanabe, Eng Siong Chng, and Haizhou Li, · 2016
Earlier work this paper cites.
“On time-frequency mask estimation for mvdr beamforming with application in robust speech recognition,”
Xiong Xiao, Shengkui Zhao, Douglas L. Jones, Eng Siong Chng, and Haizhou Li, · 2017
Earlier work this paper cites.
“Unified architecture for multichannel end-to-end speech recognition with neural beamforming,”
Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, John R. Hershey, and Xiong Xiao, · 2017
Earlier work this paper cites.
“A speech enhancement algorithm by iterating single-and multi-microphone processing and its application to robust ASR,”
Xueliang Zhang, Zhong-Qiu Wang, and DeLiang Wang, · 2017
Earlier work this paper cites.
“Multichannel signal processing with deep neural networks for automatic speech recognition,”
Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Bo Li, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew Senior, Kean Chin, et al., · 2017
Cited alongside, same era.
“Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition,”
Zhong Meng, Shinji Watanabe, John R. Hershey, and Hakan Erdogan, · 2017
Cited alongside, same era.
“Deep sets,”
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan R Salakhutdinov, and Alexander J Smola, · 2017
Cited alongside, same era.
“A simple neural network module for relational reasoning,”
Adam Santoro, David Raposo, David G. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, and Timothy Lillicrap, · 2017
Cited alongside, same era.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Cited alongside, same era.
“So-net: Self-organizing network for point cloud analysis,”
Jiaxin Li, Ben M. Chen, and Gim Hee Lee, · 2018
Later among the works it cites.
“How powerful are graph neural networks?,”
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka, · 2018
Later among the works it cites.
“gpuRIR: A python library for room impulse response simulation with gpu acceleration,”
David Diaz-Guerra, Antonio Miguel, and Jose R. Beltran, · 2018
Later among the works it cites.
“FaSNet: Low-latency adaptive beamforming for multi-microphone audio processing,”
Yi Luo, Enea Ceolini, Cong Han, Shih-Chii Liu, and Nima Mesgarani, · 2019
Closest in time.
“Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep learning based speech beamforming,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Dinei Florencio, and Mark Hasegawa-Johnson, · 2018
Cited alongside, same era.
“Performance of mask based statistical beamforming in a smart home scenario,”
Jahn Heymann, Michiel Bacchiani, and Tara N. Sainath, · 2018
Cited alongside, same era.
“Estimation of mvdr beamforming weights based on deep neural network,”
Moon Ju Jo, Geon Woo Lee, Jung Min Moon, Choongsang Cho, and Hong Kook Kim, · 2018
Cited alongside, same era.
“100 Nonspeech Sounds,” http://web.cse.ohio-state.edu/pnl/corpus/HuNonspeech/HuCorpus.html
Guoning Hu,
Cited in the paper.
Yi Luo, Zhuo Chen, and Takuya Yoshioka, · 2019
Closest in time.
“End-to-end multi-channel speech separation,”
Rongzhi Gu, Jian Wu, Shi-Xiong Zhang, Lianwu Chen, Yong Xu, Meng Yu, Dan Su, Yuexian Zou, and Dong Yu, · 2019
Closest in time.
“SDR–half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R. Hershey, · 2019
Closest in time.