Fetching the paper…
Reading the bibliography…
In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function.
H. W. Kuhn, “The hungarian method for the assignment problem,”
1955
Earlier work this paper cites.
C. Knapp and G. Carter, “The generalized correlation method for estimation of time delay,”
1976
Earlier work this paper cites.
R. Schmidt, “Multiple emitter location and signal parameter estimation,”
1986
Earlier work this paper cites.
R. Roy and T. Kailath, “ESPRIT-estimation of signal parameters via rotational invariance techniques,”
1989
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,”
1993
Earlier work this paper cites.
J.-M. Valin, F. Michaud, J. Rouat, and D. Létourneau, “Robust sound source localization using a microphone array on a mobile robot,” in
2003
Earlier work this paper cites.
R. Hartley and A. Zisserman,
2003
Earlier work this paper cites.
G. Lathoud, J.-M. Odobez, and D. Gatica-Perez, “Av16. 3: An audio-visual corpus for speaker localization and tracking,” in
2004
Earlier work this paper cites.
H. Do, H. F. Silverman, and Y. Yu, “A real-time SRP-PHAT source location implementation using stochastic region contraction (SRC) on a large-aperture microphone array,” in
2007
Earlier work this paper cites.
C. Zhang, D. Florêncio, D. E. Ba, and Z. Zhang, “Maximum likelihood sound source localization and beamforming for directional microphone arrays in distributed meetings,”
2008
Earlier work this paper cites.
F. Keyrouz, “Advanced binaural sound localization in 3-D for humanoid robots,”
2014
Earlier work this paper cites.
H.-Y. Lee, J.-W. Cho, M. Kim, and H.-M. Park, “DNN-based feature enhancement using DOA-constrained ICA for robust speech recognition,”
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in
2016
Cited alongside, same era.
M. Farmani, M. S. Pedersen, Z.-H. Tan, and J. Jensen, “Informed sound source localization using relative transfer functions for hearing aid applications,”
2017
Cited alongside, same era.
N. Wojke, A. Bewley, and D. Paulus, “Simple online and realtime tracking with a deep association metric,” in
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Z.-M. Liu, C. Zhang, and S. Y. Philip, “Direction-of-arrival estimation based on deep neural networks with robustness to array imperfections,”
2018
Cited alongside, same era.
B. Martinez, P.-C. Ma, S. Petridis, and M. Pantic, “Lipreading using temporal convolutional networks,” in
2020
Later among the works it cites.
2021
Later among the works it cites.
W. He, P. Motlicek, and J.-M. Odobez, “Neural network adaptation and data augmentation for multi-speaker direction-of-arrival estimation,”
2021
Later among the works it cites.
2021
Later among the works it cites.
H. Chen, J. Du, Y. Hu, L.-R. Dai, B.-C. Yin, and C.-H. Lee, “Correlating subword articulation with lip shapes for embedding aware audio-visual speech enhancement,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Adavanne, A. Politis, and T. Virtanen, “Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network,” in
2018
Cited alongside, same era.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in
2018
Cited alongside, same era.
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” in
2018
Cited alongside, same era.
W. He, P. Motlicek, and J.-M. Odobez, “Deep neural networks for multiple speaker detection and localization,” in
2018
Cited alongside, same era.
L. Perotin, R. Serizel, E. Vincent, and A. Guerin, “CRNN-based multiple DoA estimation using acoustic intensity features for Ambisonics recordings,”
2019
Cited alongside, same era.
Z. Tang, J. D. Kanu, K. Hogan, and D. Manocha, “Regression and classification for direction-of-arrival estimation with convolutional recurrent neural networks,” in
2019
Cited alongside, same era.
A. S. Subramanian, C. Weng, M. Yu, S.-X. Zhang, Y. Xu, S. Watanabe, and D. Yu, “Far-field location guided target speech extraction using end-to-end speech recognition objectives,” in
2020
Cited alongside, same era.
2021
Later among the works it cites.
S. Wang, A. Mesaros, T. Heittola, and T. Virtanen, “A curated dataset of urban scenes for audio-visual scene analysis,” in
2021
Later among the works it cites.
2021
Later among the works it cites.
V. Sanguineti, P. Morerio, A. Del Bue, and V. Murino, “Audio-visual localization by synthetic acoustic image generation,” in
2021
Later among the works it cites.
X. Qian, M. Madhavi, Z. Pan, J. Wang, and H. Li, “Multi-target DoA estimation with an audio-visual fusion mechanism,” in
2021
Later among the works it cites.
2022
Closest in time.
H. Chen, H. Zhou, J. Du, C.-H. Lee, J. Chen, S. Watanabe, S. M. Siniscalchi, O. Scharenborg, D.-Y. Liu, B.-C. Yin, J. Pan, J.-Q. Gao, and C. Liu, “The first multimodal information based speech processing (misp) challenge: Data, tasks, baselines and results,” in
2022
Closest in time.