Fetching the paper…
Reading the bibliography…
Multi-source localization is an important and challenging technique for multi-talker conversation analysis.
1910
Earlier work this paper cites.
J. B. Allen, D. A. Berkley, Image method for efficiently simulating small-room acoustics, The Journal of the Acoustical Society of America 65 (4) (1979) 943–950
1979
Earlier work this paper cites.
R. Schmidt, Multiple emitter location and signal parameter estimation, IEEE transactions on antennas and propagation 34 (3) (1986) 276–280
1986
Earlier work this paper cites.
R. Roy, T. Kailath, ESPRIT-estimation of signal parameters via rotational invariance techniques, IEEE Transactions on Acoustics, Speech, and Signal Processing 37 (7) (1989) 984–995
1989
Earlier work this paper cites.
Yeo-Sun Yoon, L. M. Kaplan, J. H. McClellan, TOPS: new DOA estimator for wideband signals, IEEE Transactions on Signal Processing 54 (6) (2006) 1977–1989
1989
Earlier work this paper cites.
T. Robinson, J. Fransen, D. Pye, J. Foote, S. Renals, WSJCAMO: a British English speech corpus for large vocabulary continuous speech recognition, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Vol. 1, 1995, pp. 81–84
1995
Earlier work this paper cites.
K. Nakadai, K. Hidai, H. Mizoguchi, H. Okuno, H. Kitano, Real-time auditory and visual multiple-object tracking for humanoids, in: International Joint Conferences on Artificial Intelligence (IJCAI), 2001, pp. 1425–1432
2001
Earlier work this paper cites.
E. D. Di Claudio, R. Parisi, WAVES: weighted average of signal subspaces for robust wideband direction finding, IEEE Transactions on Signal Processing 49 (10) (2001) 2179–2191
2001
Earlier work this paper cites.
E. Levina, P. Bickel, The earth mover’s distance is the mallows distance: some insights from statistics, in: IEEE International Conference on Computer Vision (ICCV), Vol. 2, 2001, pp. 251–256
2001
Earlier work this paper cites.
M. Lincoln, I. McCowan, J. Vepa, H. K. Maganti, The multi-channel wall street journal audio visual corpus (MC-WSJ-AV): specification and initial experiments, in: IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), 2005, pp. 357–362
2005
Earlier work this paper cites.
K. Nakadai, T. Takahashi, H. Okuno, H. Nakajima, Y. Hasegawa, H. Tsujino, Design and implementation of robot audition system ‘HARK’ - open source software for listening to three simultaneous speakers, Advanced Robotics 24 (5-6) (2010) 739–761
2010
Earlier work this paper cites.
M. Souden, J. Benesty, S. Affes, On optimal frequency-domain multichannel linear filtering for noise reduction, IEEE Transactions on Audio, Speech, and Language Processing 18 (2) (2010) 260–276
2010
Earlier work this paper cites.
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, B. Juang, Speech dereverberation based on variance-normalized delayed linear prediction, IEEE Transactions on Audio, Speech, and Language Processing 18 (7) (2010) 1717–1731
2010
Earlier work this paper cites.
A. Narayanan, D. Wang, Ideal ratio mask estimation using deep neural networks for robust speech recognition, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2013, pp. 7092–7096
2013
Earlier work this paper cites.
Y. Wang, A. Narayanan, D. Wang, On training targets for supervised speech separation, IEEE/ACM Transactions on Audio, Speech, and Language Processing 22 (12) (2014) 1849–1858
2014
Earlier work this paper cites.
D. Salvati, C. Drioli, G. L. Foresti, Incoherent frequency fusion for broadband steered response power algorithms in noisy environments, IEEE Signal Processing Letters 21 (5) (2014) 581–585
2014
Earlier work this paper cites.
T. Hirvonen, Classification of spatial audio location and content using convolutional neural networks, in: 138th Audio Engineering Society Convention, Vol. 2, 2015
2015
Earlier work this paper cites.
D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: International Conference on Learning Representations (ICLR), 2015
2015
Earlier work this paper cites.
R. Takeda, K. Komatani, Sound source localization based on deep neural networks with directional activate function exploiting phase information, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2016, pp. 405–409
2016
Earlier work this paper cites.
F. Vesperini, P. Vecchiotti, E. Principi, S. Squartini, F. Piazza, A neural network based algorithm for speaker localization in a multi-room environment, in: 26th IEEE International Workshop on Machine Learning for Signal Processing (MLSP), 2016, pp. 1–6
2016
Cited alongside, same era.
K. Veselý, S. Watanabe, K. Žmolíková, M. Karafiát, L. Burget, J. H. Černocký, Sequence summarizing neural network for speaker adaptation, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2016, pp. 5315–5319
2016
Cited alongside, same era.
N. Yalta, K. Nakadai, T. Ogata, Sound source localization using deep learning models, Journal of Robotics and Mechatronics 29 (2017) 37–48
2017
Cited alongside, same era.
S. Chakrabarty, E. A. Habets, Broadband DOA estimation using convolutional neural networks trained with noise signals, in: IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2017, pp. 136–140
2017
Cited alongside, same era.
R. Scheibler, E. Bezzam, I. Dokmanić, Pyroomacoustics: A python package for audio room simulation and array processing algorithms, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2018, pp. 351–355
2018
Later among the works it cites.
K. Kinoshita, L. Drude, M. Delcroix, T. Nakatani, Listening to each speaker one by one with recurrent selective hearing networks, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2018, pp. 5064–5068
2018
Later among the works it cites.
T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dimitriadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang, A. Hurvitz, L. Jiang, S. Koubi, E. Krupka, I. Leichter, C. Liu, P. Parthasarathy, A. Vinnikov, L. Wu, X. Xiao, W. Xiong, H. Wang, Z. Wang, J. Zhang, Y. Zhao, T. Zhou, Advances in online audio-visual meeting transcription, in: IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, pp. 276–283
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Yu, M. Kolbæk, Z. Tan, J. Jensen, Permutation invariant training of deep models for speaker-independent multi-talker speech separation, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2017, pp. 241–245
2017
Cited alongside, same era.
S. Kim, T. Hori, S. Watanabe, Joint CTC-attention based end-to-end speech recognition using multi-task learning, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2017, pp. 4835–4839
2017
Cited alongside, same era.
J. Barker, S. Watanabe, E. Vincent, J. Trmal, The fifth CHiME speech separation and recognition challenge: Dataset, task and baselines, in: Interspeech, 2018, pp. 1561–1565
2018
Cited alongside, same era.
S. Adavanne, A. Politis, J. Nikunen, T. Virtanen, Sound event localization and detection of overlapping sources using convolutional recurrent neural networks, IEEE Journal of Selected Topics in Signal Processing 13 (1) (2018) 34–48
2018
Cited alongside, same era.
S. Sivasankaran, E. Vincent, D. Fohr, Keyword based speaker localization: Localizing a target speaker in a multi-speaker environment, in: Interspeech, 2018, pp. 2703–2707
2018
Cited alongside, same era.
M. Delcroix, K. Zmolikova, K. Kinoshita, A. Ogawa, T. Nakatani, Single channel target speaker extraction and recognition with speaker beam, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018, pp. 5554–5558
2018
Cited alongside, same era.
H. W. Löllmann, C. Evers, A. Schmidt, H. Mellmann, H. Barfuss, P. A. Naylor, W. Kellermann, The locata Challenge data corpus for acoustic source localization and tracking, in: IEEE 10th Sensor Array and Multichannel Signal Processing Workshop, 2018, pp. 410–414
2018
Cited alongside, same era.
S. Adavanne, A. Politis, T. Virtanen, Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network, in: European Signal Processing Conference (EUSIPCO), 2018, pp. 1462–1466
2018
Cited alongside, same era.
R. Haeb-Umbach, S. Watanabe, T. Nakatani, M. Bacchiani, B. Hoffmeister, M. L. Seltzer, H. Zen, M. Souden, Speech processing for digital home assistants: Combining signal processing with deep-learning techniques, IEEE Signal processing magazine 36 (6) (2019) 111–124
2019
Later among the works it cites.
S. Chakrabarty, E. A. Habets, Multi-speaker DOA estimation using deep convolutional networks trained with noise signals, IEEE Journal of Selected Topics in Signal Processing 13 (1) (2019) 8–21
2019
Later among the works it cites.
L. Perotin, A. Défossez, E. Vincent, R. Serizel, A. Guérin, Regression versus classification for neural network based audio source localization, in: IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2019, pp. 343–347
2019
Later among the works it cites.
W. Zhang, Y. Zhou, Y. Qian, Robust DOA estimation based on convolutional neural network and time-frequency masking, in: Interspeech, 2019, pp. 2703–2707
2019
Later among the works it cites.
Z.-Q. Wang, X. Zhang, D. Wang, Robust speaker localization guided by deep learning-based time-frequency masking, IEEE/ACM Transactions on Audio, Speech, and Language Processing 27 (1) (2019) 178–188
2019
Later among the works it cites.
X. Chang, W. Zhang, Y. Qian, J. Le Roux, S. Watanabe, MIMO-Speech: End-to-end multi-channel multi-speaker speech recognition, in: IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, pp. 237–244
2019
Later among the works it cites.
F. Bahmaninezhad, J. Wu, R. Gu, S.-X. Zhang, Y. Xu, M. Yu, D. Yu, A comprehensive study of speech separation: Spectrogram vs waveform separation, in: Interspeech, 2019, pp. 4574–4578
2019
Later among the works it cites.
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. E. Y. Soplin, R. Yamamoto, X. Wang, S. Watanabe, T. Yoshimura, W. Zhang, A comparative study on Transformer vs RNN in speech applications, in: IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, pp. 449–456
2019
Later among the works it cites.
W. Mack, U. Bharadwaj, S. Chakrabarty, E. A. P. Habets, Signal-aware broadband DOA estimation using attention mechanisms, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 4930–4934
2020
Later among the works it cites.
A. S. Subramanian, C. Weng, M. Yu, S. Zhang, Y. Xu, S. Watanabe, D. Yu, Far-field location guided target speech extraction using end-to-end speech recognition objectives, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2020, pp. 7299–7303
2020
Later among the works it cites.
X. Chang, W. Zhang, Y. Qian, J. Le Roux, S. Watanabe, End-to-end multi-speaker speech recognition with transformer, in: IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2020, pp. 6134–6138
2020
Later among the works it cites.
J. Shi, X. Chang, P. Guo, S. Watanabe, Y. Fujita, J. Xu, B. Xu, L. Xie, Sequence to multi-sequence learning via conditional chain mapping for mixture signals, in: Advances in Neural Information Processing Systems, 2020, pp. 3735–3747
2020
Later among the works it cites.
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, D. Snyder, A. S. Subramanian, J. Trmal, B. B. Yair, C. Boeddeker, Z. Ni, Y. Fujita, S. Horiguchi, N. Kanda, T. Yoshioka, N. Ryant, CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings, in: 6th International Workshop on Speech Processing in Everyday Environments, 2020
2020
Later among the works it cites.
K. Shimada, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, ACCDOA: Activity-coupled cartesian direction of arrival representation for sound event localization and detection, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 915–919
2021
Closest in time.
A. S. Subramanian, C. Weng, S. Watanabe, M. Yu, Y. Xu, S.-X. Zhang, D. Yu, Directional ASR: A new paradigm for e2e multi-speaker speech recognition with source localization, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 8433–8437
2021
Closest in time.