Fetching the paper…
Reading the bibliography…
Beamforming has been extensively investigated for multi-channel audio processing tasks.
“High-resolution frequency-wavenumber spectrum analysis,”
Jack Capon, · 1969
Earlier work this paper cites.
“The generalized correlation method for estimation of time delay,”
Charles Knapp and Glifford Carter, · 1976
Earlier work this paper cites.
“Image method for efficiently simulating small-room acoustics,”
Jont B Allen and David A Berkley, · 1979
Earlier work this paper cites.
“Neural network adaptive beamforming for robust multichannel speech recognition.,”
Bo Li, Tara N Sainath, Ron J Weiss, Kevin W Wilson, and Michiel Bacchiani, · 1980
Earlier work this paper cites.
“Improved mvdr beamforming using single-channel mask prediction networks.,”
Hakan Erdogan, John R Hershey, Shinji Watanabe, Michael I Mandel, and Jonathan Le Roux, · 1985
Earlier work this paper cites.
“Timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,”
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, and David S Pallett, · 1993
Earlier work this paper cites.
“A robust method for speech signal time-delay estimation in reverberant rooms,”
Michael S Brandstein and Harvey F Silverman, · 1997
Earlier work this paper cites.
“Performance of time-and frequency-domain binaural beamformers based on recorded signals from real rooms,”
Michael E Lockwood, Douglas L Jones, Robert C Bilger, Charissa R Lansing, William D O’Brien Jr, Bruce C Wheeler, and Albert S Feng, · 2004
Earlier work this paper cites.
“Blind acoustic beamforming based on generalized eigenvalue decomposition,”
Ernst Warsitz and Reinhold Haeb-Umbach, · 2007
Earlier work this paper cites.
“Acoustic beamforming for hearing aid applications,”
Simon Doclo, Sharon Gannot, Marc Moonen, and Ann Spriet, · 2010
Earlier work this paper cites.
“A fast signal subspace approach for the determination of absolute levels from phased microphone array measurements,”
Ennes Sarradj, · 2010
Earlier work this paper cites.
Time-Domain MVDR Array Filter for Speech Enhancement
Mingsian Bai, Jeong-Guon Ih, and Jacob Benesty, · 2013
Earlier work this paper cites.
“Performance comparison of time-domain and frequency-domain beamforming techniques for sensor array processing,”
Umar Hamid, Rahim Ali Qamar, and Kashif Waqas, · 2014
Earlier work this paper cites.
“Speaker location and microphone spacing invariant acoustic modeling from raw multichannel waveforms,”
Tara N Sainath, Ron J Weiss, Kevin W Wilson, Arun Narayanan, Michiel Bacchiani, et al., · 2015
Earlier work this paper cites.
“Blstm supported gev beamformer front-end for the 3rd chime challenge,”
Jahn Heymann, Lukas Drude, Aleksej Chinaev, and Reinhold Haeb-Umbach, · 2015
Earlier work this paper cites.
“The third ‘chime’ speech separation and recognition challenge: Dataset, task and baselines,”
Jon Barker, Ricard Marxer, Emmanuel Vincent, and Shinji Watanabe, · 2015
Cited alongside, same era.
“A study of learning based beamforming methods for speech recognition,”
Xiong Xiao, Chenglin Xu, Zhaofeng Zhang, Shengkui Zhao, Sining Sun, Shinji Watanabe, Longbiao Wang, Lei Xie, Douglas L Jones, Eng Siong Chng, et al., · 2016
Cited alongside, same era.
“Deep beamforming networks for multi-channel speech recognition,”
Xiong Xiao, Shinji Watanabe, Hakan Erdogan, Liang Lu, John Hershey, Michael L Seltzer, Guoguo Chen, Yu Zhang, Michael Mandel, and Dong Yu, · 2016
Cited alongside, same era.
“Beamforming networks using spatial covariance features for far-field speech recognition,”
Xiong Xiao, Shinji Watanabe, Eng Siong Chng, and Haizhou Li, · 2016
Cited alongside, same era.
“Neural network based spectral mask estimation for acoustic beamforming,”
Jahn Heymann, Lukas Drude, and Reinhold Haeb-Umbach, · 2016
Cited alongside, same era.
“A speech enhancement algorithm by iterating single-and multi-microphone processing and its application to robust asr,”
Xueliang Zhang, Zhong-Qiu Wang, and DeLiang Wang, · 2017
Later among the works it cites.
“Permutation invariant training of deep models for speaker-independent multi-talker speech separation,”
D. Yu, M. Kolbæk, Z. Tan, and J. Jensen, · 2017
Later among the works it cites.
“Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,”
Morten Kolbæk, Dong Yu, Zheng-Hua Tan, Jesper Jensen, Morten Kolbaek, Dong Yu, Zheng-Hua Tan, and Jesper Jensen, · 2017
Later among the works it cites.
“Estimation of mvdr beamforming weights based on deep neural network,”
Moon Ju Jo, Geon Woo Lee, Jung Min Moon, Choongsang Cho, and Hong Kook Kim, · 2018
Later among the works it cites.
“Exploring practical aspects of neural mask-based beamforming for far-field speech recognition,”
Christoph Boeddeker, Hakan Erdogan, Takuya Yoshioka, and Reinhold Haeb-Umbach, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Multi-channel speech recognition: Lstms all the way through,”
Hakan Erdogan, Tomoki Hayashi, John R Hershey, Takaaki Hori, Chiori Hori, Wei-Ning Hsu, Suyoun Kim, Jonathan Le Roux, Zhong Meng, and Shinji Watanabe, · 2016
Cited alongside, same era.
“A consolidated perspective on multimicrophone speech enhancement and source separation,”
Sharon Gannot, Emmanuel Vincent, Shmulik Markovich-Golan, and Alexey Ozerov, · 2017
Cited alongside, same era.
“Multichannel signal processing with deep neural networks for automatic speech recognition,”
Tara N Sainath, Ron J Weiss, Kevin W Wilson, Bo Li, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew Senior, Kean Chin, et al., · 2017
Cited alongside, same era.
Zhong Meng, Shinji Watanabe, John R Hershey, and Hakan Erdogan, · 2017
Cited alongside, same era.
“On time-frequency mask estimation for mvdr beamforming with application in robust speech recognition,”
Xiong Xiao, Shengkui Zhao, Douglas L Jones, Eng Siong Chng, and Haizhou Li, · 2017
Cited alongside, same era.
“Multichannel end-to-end speech recognition,”
Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, and John R Hershey, · 2017
Cited alongside, same era.
“Unified architecture for multichannel end-to-end speech recognition with neural beamforming,”
Tsubasa Ochiai, Shinji Watanabe, Takaaki Hori, John R Hershey, and Xiong Xiao, · 2017
Cited alongside, same era.
Later among the works it cites.
“Online integration of dnn-based and spatial clustering-based mask estimation for robust mvdr beamforming,”
Yutaro Matsui, Tomohiro Nakatani, Marc Delcroix, Keisuke Kinoshita, Nobutaka Ito, Shoko Araki, and Shoji Makino, · 2018
Later among the works it cites.
“Performance of mask based statistical beamforming in a smart home scenario,”
Jahn Heymann, Michiel Bacchiani, and Tara N Sainath, · 2018
Later among the works it cites.
“Wave-u-net: A multi-scale neural network for end-to-end audio source separation,”
Daniel Stoller, Sebastian Ewert, and Simon Dixon, · 2018
Later among the works it cites.
“Raw multi-channel audio source separation using multi-resolution convolutional auto-encoders,”
Emad M Grais, Dominic Ward, and Mark D Plumbley, · 2018
Later among the works it cites.
“Deep learning based speech beamforming,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, Dinei Florencio, and Mark Hasegawa-Johnson, · 2018
Later among the works it cites.
“Tasnet: Surpassing ideal time-frequency masking for speech separation,”
Yi Luo and Nima Mesgarani, · 2018
Later among the works it cites.
“Speaker-independent speech separation with deep attractor network,”
Yi Luo, Zhuo Chen, and Nima Mesgarani, · 2018
Later among the works it cites.
“gpurir: A python library for room impulse response simulation with gpu acceleration,”
David Diaz-Guerra, Antonio Miguel, and Jose R Beltran, · 2018
Later among the works it cites.
“Sdr – half-baked or well done?,”
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R. Hershey, · 2019
Closest in time.