Fetching the paper…
Reading the bibliography…
In this paper, we propose a convolutional recurrent neural network for joint sound event localization and detection (SELD) of multiple overlapping sound events in three-dimensional (3D) space.
J. B. Allen and D. A. Berkley, “Image method for efficiently simulating small-room acoustics,” in The Journal of the Acoustical Society of America , vol. 65, no. 4, 1979
1979
Earlier work this paper cites.
R. O. Schmidt, “Multiple emitter location and signal parameter estimation,” in IEEE Transactions on Antennas and Propagation , vol. 34, no. 3, 1986
1986
Earlier work this paper cites.
R. Roy and T. Kailath, “ESPRIT-estimation of signal parameters via rotational invariance techniques,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 37, no. 7, 1989
1989
Earlier work this paper cites.
B. Ottersten, M. Viberg, P. Stoica, and A. Nehorai, “Exact and large sample maximum likelihood techniques for parameter estimation and detection in array processing,” in Radar Array Processing. Springer Series in Information Sciences , 1993
1993
Earlier work this paper cites.
H. Wang and P. Chu, “Voice source localization for automatic camera pointing system in videoconferencing,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 1997
1997
Earlier work this paper cites.
M. S. Brandstein and H. F. Silverman, “A high-accuracy, low-latency technique for talker localization in reverberant environments using microphone arrays,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 1997
1997
Earlier work this paper cites.
Y. Huang, J. Benesty, G. Elko, and R. Mersereati, “Real-time passive source localization: a practical linear-correction least-squares approach,” in IEEE Transactions on Speech and Audio Processing , vol. 9, no. 8, 2001
2001
Earlier work this paper cites.
J. H. DiBiase, H. F. Silverman, and M. S. Brandstein, “Robust localization in reverberant rooms,” in Microphone Arrays . Springer, 2001
2001
Earlier work this paper cites.
C. Busso, S. Hernanz, C.-W. Chu, S.-i. Kwon, S. Lee, P. G. Georgiou, I. Cohen, and S. Narayanan, “Smart room: participant and speaker localization and identification,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2005
2005
Earlier work this paper cites.
H. Teutsch, Modal array signal processing: principles and applications of acoustic wavefield decomposition . Springer, 2007, vol. 348
2007
Earlier work this paper cites.
J.-M. Valin, F. Michaud, and J. Rouat, “Robust localization and tracking of simultaneous moving sound sources using beamforming and particle filtering,” Robotics and Autonomous Systems , vol. 55, no. 3, pp. 216–228, 2007
2007
Earlier work this paper cites.
A. Temko, C. Nadeu, and J.-I. Biel, “Acoustic event detection: SVM-based system and evaluation setup in CLEAR’07,” in Multimodal Technologies for Perception of Humans . Springer, 2008
2008
Earlier work this paper cites.
M. Wölfel and J. McDonough, Distant speech recognition . John Wiley & Sons, 2009
2009
Earlier work this paper cites.
S. Chu, S. Narayanan, and C. J. Kuo, “Environmental sound recognition with time-frequency audio features,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 17, no. 6, 2009
2009
Earlier work this paper cites.
G. Enzner, “3D-continuous-azimuth acquisition of head-related impulse responses using multi-channel adaptive filtering,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2009
2009
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Eronen, and T. Virtanen, “Acoustic event detection in real-life recordings,” in European Signal Processing Conference (EUSIPCO) , 2010
2010
Earlier work this paper cites.
T. Butko, F. G. Pla, C. Segura, C. Nadeu, and J. Hernando, “Two-source acoustic event detection and localization: Online implementation in a smart-room,” in European Signal Processing Conference (EUSIPCO) , 2011
2011
Earlier work this paper cites.
T. A. Marques et al. , “Estimating animal population density using passive acoustics,” in Biological reviews of the Cambridge Philosophical Society , vol. 88, no. 2, 2012
2012
Earlier work this paper cites.
D. Khaykin and B. Rafaely, “Acoustic analysis by spherical microphone array processing of room impulse responses,” The Journal of the Acoustical Society of America , vol. 132, no. 1, 2012
2012
Earlier work this paper cites.
P. Swietojanski, A. Ghoshal, and S. Renals, “Convolutional neural networks for distant speech recognition,” in IEEE Signal Processing Letters , vol. 21, 2014
2014
Earlier work this paper cites.
B. J. Furnas and R. L. Callas, “Using automated recorders and occupancy models to monitor common forest birds across a large geographic region,” in Journal of Wildlife Management , vol. 79, no. 2, 2014
2014
Earlier work this paper cites.
J. Traa and P. Smaragdis, “Multiple speaker tracking with the Factorial Von Mises-Fisher filter,” in IEEE International Workshop on Machine Learning for Signal Processing (MLSP) , 2014
2014
Earlier work this paper cites.
R. Chakraborty and C. Nadeu, “Sound-model-based acoustic source localization using distributed microphone arrays,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in ACM International Conference on Multimedia (ACM-MM) , 2014
2014
Cited alongside, same era.
P. Foggia, N. Petkov, A. Saggese, N. Strisciuglio, and M. Vento, “Audio surveillance of roads: A system for detecting anomalous sounds,” in IEEE Transactions on Intelligent Transportation Systems , vol. 17, no. 1, 2015
2015
Cited alongside, same era.
X. Xiao, S. Zhao, X. Zhong, D. L. Jones, E. S. Chng, and H. Li, “A learning-based approach to direction of arrival estimation in noisy and reverberant environments,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Cited alongside, same era.
T. Hirvonen, “Classification of spatial audio location and content using convolutional neural networks,” in Audio Engineering Society Convention 138 , 2015
2015
Cited alongside, same era.
P. W. Wessels, J. V. Sande, and F. V. der Eerden, “Detection and localization of impulsive sound events for environmental noise assessment,” in The Journal of the Acoustical Society of America 141 , vol. 141, no. 5, 2017
2017
Later among the works it cites.
S. Chakrabarty and E. A. P. Habets, “Broadband DOA estimation using convolutional neural networks trained with noise signals,” in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2017
2017
Later among the works it cites.
——, “Multi-speaker localization using convolutional neural network trained with noise,” in Neural Information Processing Systems (NIPS) , 2017
2017
Later among the works it cites.
M. Yiwere and E. J. Rhee, “Distance estimation and localization of sound sources in reverberant conditions using deep neural networks,” in International Journal of Applied Engineering Research , vol. 12, no. 22, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Roden, N. Moritz, S. Gerlach, S. Weinzierl, and S. Goetze, “On sound source localization of speech signals using deep neural networks,” in Deutsche Jahrestagung für Akustik (DAGA) , 2015
2015
Cited alongside, same era.
E. Çakır, T. Heittola, H. Huttunen, and T. Virtanen, “Polyphonic sound event detection using multi-label deep neural networks,” in IEEE International Joint Conference on Neural Networks (IJCNN) , 2015
2015
Cited alongside, same era.
H. Zhang, I. McLoughlin, and Y. Song, “Robust sound event recognition using convolutional neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” International Conference on Machine Learning , 2015
2015
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third CHiME speech separation and recognition challenge: Dataset, task and baselines,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015
2015
Cited alongside, same era.
R. Takeda and K. Komatani, “Sound source localization based on deep neural networks with directional activate function exploiting phase information,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Cited alongside, same era.
——, “Discriminative multiple sound source localization based on deep neural networks using independent location model,” in IEEE Spoken Language Technology Workshop (SLT) , 2016
2016
Cited alongside, same era.
Y. Sun, J. Chen, C. Yuen, and S. Rahardja, “Indoor sound source localization with probabilisitic neural network,” in IEEE Transactions on Industrial Electronics , vol. 29, no. 1, 2017
2017
Later among the works it cites.
T. Hayashi, S. Watanabe, T. Toda, T. Hori, J. L. Roux, and K. Takeda, “Duration-controlled LSTM for polyphonic sound event detection,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 11, 2017
2017
Later among the works it cites.
M. Zöhrer and F. Pernkopf, “Virtual adversarial training and data augmentation for acoustic event detection with gated recurrent neural networks,” in INTERSPEECH , 2017
2017
Later among the works it cites.
H. Lim, J. Park, K. Lee, and Y. Han, “Rare sound event detection using 1D convolutional recurrent neural networks,” in Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017
2017
Later among the works it cites.
E. Çakır, G. Parascandolo, T. Heittola, H. Huttunen, and T. Virtanen, “Convolutional recurrent neural networks for polyphonic sound event detection,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 6, 2017
2017
Later among the works it cites.
S. Adavanne and T. Virtanen, “A report on sound event detection with different binaural features,” in Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017
2017
Later among the works it cites.
S. Adavanne, P. Pertilä, and T. Virtanen, “Sound event detection using spatial features and convolutional recurrent neural network,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017
2017
Later among the works it cites.
I.-Y. Jeong, S. Lee1, Y. Han, and K. Lee, “Audio event detection using multiple-input convolutional neural network,” in Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017
2017
Later among the works it cites.
J. Zhou, “Sound event detection in multichannel audio LSTM network,” in Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017
2017
Later among the works it cites.
R. Lu and Z. Duan, “Bidirectional gru for sound event detection,” in Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017
2017
Later among the works it cites.
A. Mesaros, T. Heittola, A. Diment, B. Elizalde, A. Shah, E. Vincent, B. Raj, and T. Virtanen, “DCASE 2017 challenge setup: tasks, datasets and baseline system,” in Detection and Classification of Acoustic Scenes and Events 2017 Workshop (DCASE) , 2017
2017
Later among the works it cites.
W. He, P. Motlicek, and J.-M. Odobez, “Deep neural networks for multiple speaker detection and localization,” in International Conference on Robotics and Automation (ICRA) , 2018
2018
Closest in time.
E. L. Ferguson, S. B. Williams, and C. T. Jin, “Sound source localization in a multipath environment using convolutional neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018
2018
Closest in time.
S. Adavanne, A. Politis, and T. Virtanen, “Direction of arrival estimation for multiple sound sources using convolutional recurrent neural network,” in European Signal Processing Conference (EUSIPCO) , 2018
2018
Closest in time.
S. Adavanne, A. Politis, and T. Virtanen, “Multichannel sound event detection using 3D convolutional neural networks for learning inter-channel features,” in IEEE International Joint Conference on Neural Networks (IJCNN) , 2018
2018
Closest in time.
F. Chollet, “Keras v2.0.8,” 2015, accessed on 7 May 2018. [Online]. Available: https://github.com/fchollet/keras
2018
Closest in time.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, accessed on 7 May 2018. [Online]. Available: https://www.tensorflow.org/
2018
Closest in time.
E. Benetos, M. Lagrange, and G. Lafay, “Sound event detection in synthetic audio,” 2016, accessed on 7 May 2018. [Online]. Available: https://archive.org/details/dcase2016_task2_train_dev
2018
Closest in time.