Fetching the paper…
Reading the bibliography…
A major advantage of a deep convolutional neural network (CNN) is that the focused receptive field size is increased by stacking multiple convolutional layers.
1912
Earlier work this paper cites.
D. Rumelhart, G. Hinton, R. Williams, “Learning representations by back-propagating errors,” Nature , 323 (6088), 533-536, 1986
1986
Earlier work this paper cites.
S. Hochreiter, J. Schmidhuber, “Long Short-Term Memory,” Neural Computation , 9 (8), pp. 1735-1780, 1997
1997
Earlier work this paper cites.
O. Maron, A. L. Ratan, “Multiple-Instance Learning for Natural Scene Classification,” in Proc. 15th Int. Conf. on Machine Learning , pp. 341-349, 1998
1998
Earlier work this paper cites.
J. Ramon, L. De Raedt, “Multi instance neural networks,” in Proc. of the ICML-2000 workshop on attribute-value and relational learning , pp. 53-60, 2000
2000
Earlier work this paper cites.
A. Temko, C. Nadeu, and J. I. Biel, “Acoustic Event Detection: SVM-Based System and Evaluation Setup in CLEAR’07,” Multimodel technologies for perception of humans , pp. 354-363, 2008
2008
Earlier work this paper cites.
C. Zieger, “An HMM based system for acoustic event detection,” Multimodel technologies for perception of humans , pp. 338-344, 2008
2008
Earlier work this paper cites.
X. Zhuang, X. Zhou, M. A. Hasegawa-johnson, T. S. Huang, “Real-world acoustic event detection,” Pattern Recognition Letters , vol. 31, no. 12, pp. 1543-1551, 2010
2010
Earlier work this paper cites.
C. V. Cotton, D. P. W. Ellis, “Spectral vs. spectro-temporal features for acoustic event detection,” in Proc. of IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , pp. 69-72, 2010
2010
Earlier work this paper cites.
D. Giannoulisy, E. Benetosx, D. Stowelly, M. Rossignolz, M. Lagrangez and M. Plumbley, “Detection and Classification of Acoustic Scenes and Events: an IEEE AASP Challenge,” IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , 2013
2013
Earlier work this paper cites.
T. Heittola, A. Mesaros, A. Eronen, and T. Virtanen, “Context-dependent sound event detection,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 1, pp. 1-13, 2013
2013
Earlier work this paper cites.
Z. Huang, Y. Cheng, K. Li, V. Hautamaki, C. Lee, “A Blind Segmentation Approach to Acoustic Event Detection Based on I-Vector,” in Proc. of the Annual Conference of the International Speech Communication Association (Interspeech) , pp. 2282-2286, 2013
2013
Earlier work this paper cites.
A. Plinge, R. Grzeszick, and G. Fink, “A Bag-of-Features approach to acoustic event detection,” in Proc. of IEEE Int. Conf. on Acoustics, Speech, and Signal Process (ICASSP) ., pp. 3732-3736, 2014
2014
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, C. Hori, “Sparse representation based on a bag of spectral exemplars for acoustic event detection,” in Proc. of IEEE Int. Conf. on Acoustics, Speech, and Signal Process (ICASSP) , pp. 6255-6259, 2014
2014
Earlier work this paper cites.
K. Cho, B. Merrienboer, D. Bahdanau, Y. Bengio, “On the Properties of Neural Machine Translation: Encoder-Decoder Approaches,” the 8-th Workshop on Syntax, Semantics and Structure in Statistical Translation , SSST-8, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Salamon, C. Jacoby and J. P. Bello, “A Dataset and Taxonomy for Urban Sound Research,” the 22nd ACM International Conference on Multimedia , Orlando USA, Nov. 2014
2014
Earlier work this paper cites.
Diederik P. Kingma, Jimmy Ba, “Adam: A Method for Stochastic Optimization,” the 3rd International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
Karol J. Piczak, “Environmental sound classification with convolutional neural networks,” IEEE 25th International Workshop on Machine Learning for Signal Processing (MLSP) , 12 November 2015
2015
Cited alongside, same era.
M. Luong, H. Pham, C. Manning, “Effective Approaches to Attention-based Neural Machine Translation,” in Proc. of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , pp. 1412-1421, 2015
2015
Cited alongside, same era.
S. Ioffe, C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” In proceedings of the 32nd International Conference on International Conference on Machine Learning (ICML) , pp. 448-456, 2015
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Lee, S. Lee, Y. Han, K. Lee, “Ensemble of Convolutional Neural Networks for Weakly-Supervised Sound Event Detection using Multiple Scale Input,” Detection and Classification of Acoustic Scenes and Events (DCASE) , 2017
2017
Later among the works it cites.
Q. Kong, Y. Xu, W. Wang, MD. Plumbley, “A joint detection-classification model for audio tagging of weakly labelled data,” in Proc. of IEEE Int. Conf. on Acoustics, Speech, and Signal Process (ICASSP) , pp. 641-645, 2017
2017
Later among the works it cites.
Y. Xu, Q. Kong, Q. Huang, W. Wang, MD. Plumbley, “Attention and Localization based on a Deep Convolutional Recurrent Model for Weakly Supervised Audio Tagging,” in Proc. of the Annual Conference of the International Speech Communication Association (Interspeech) , pp. 3083-3087, 2017
2017
Later among the works it cites.
T. W. Su, J. Y. Liu, Y. H. Yang, “Weakly-supervised audio event detection using event-specific Gaussian filters and fully convolutional networks,” in Proc. of IEEE Int. Conf. on Acoustics, Speech, and Signal Process (ICASSP) , pp. 791-795, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Gorin, N. Makhazhanov, and N. Shmyrev, “DCASE 2016 sound event detection system based on convolutional neural network,” Tech. Rep., DCASE2016 Challenge , 2016
2016
Cited alongside, same era.
A. Kumar, B. Raj, “Audio Event Detection using Weakly Labeled Data,” in Proc. of the ACM on Multimedia Conference , pp. 1038-1047, 2016
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. of IEEE Int. Conf. on Acoustics, Speech, and Signal Process (ICASSP) , pp. 4960-4964, 2016
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, J. Sun, “Deep Residual Learning for Image Recognition,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, J. Sun, “Identity Mappings in Deep Residual Networks,” European Conference on Computer Vision , 2016
2016
Cited alongside, same era.
J. Y. Liu, Y. H. Yang, “Event Localization in Music Auto-tagging,” in Proc. of ACM on Multimedia Conference , pp. 1048-1057, 2016
2016
Cited alongside, same era.
A. Veit, M Wilber, S. Belongie, “Residual Networks Behave Like Ensembles of Relatively Shallow Networks,” Advances in neural information processing systems , 2016
2016
Cited alongside, same era.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger, “Deep Networks with Stochastic Depth,” European Conference on Computer Vision , 2016
2016
Cited alongside, same era.
2017
Later among the works it cites.
A. Kumar, B. Raj, “Deep CNN Framework for Audio Event Recognition using Weakly Labeled Web Data,” in NIPS Workshop on Machine Learning for Audio , 2017
2017
Later among the works it cites.
J. Gemmeke, D. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events, ICASSP , 2017
2017
Later among the works it cites.
A. Mesaros, T. Heittola, A. Diment, B. Elizalde, A. Shah, E. Vincent, B. Raj, and T. Virtanen, “Dcase 2017 challenge setup: Tasks, datasets and baseline system,” DCASE Workshop , 2017
2017
Later among the works it cites.
F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017
2017
Later among the works it cites.
Y. Guo, M. Xu, J. Wu, Y. Wang, K. Hoashi, “Multi-scale convolutional recurrent neural network with ensemble method for weakly labeled sound event detection,” Detection and Classification of Acoustic Scenes and Events (DCASE) , 2018
2018
Later among the works it cites.
Y. Xu, Q. Kong, W. Wang, MD. Plumbley, “Large-scale weakly supervised audio classification using gated convolutional neural network,” ICASSP , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
X. Lu, P. Shen, S. Li, Y. Tsao, H. Kawai, “Temporal Attentive Pooling for Acoustic Event Detection,” in Proc. of the Annual Conference of the International Speech Communication Association (Interspeech) , pp. 1354-1357, 2018
2018
Later among the works it cites.
S. Chou, J. Jang, Y. Yang, “Learning to Recognize Transient Sound Events using Attentional Supervision,” in Proc. Int. Joint Conf. Artificial Intelligence (IJCAI) , 2018
2018
Later among the works it cites.
X. Wang, Y. Yan, P. Tang, X. Bai, W. Liu, “Revisiting Multiple Instance Neural Networks,” Pattern Recognition , No. 74, pp. 15-24, 2018
2018
Later among the works it cites.
S. Tseng, J. Li, Y. Wang, F. Metze, J. Szurley, S. Das, “Multiple instance deep learning for weakly supervised small-footprint audio event detection,” in Proc. of the Annual Conference of the International Speech Communication Association (Interspeech) , pp. 3279-3283, Sep. 2018
2018
Later among the works it cites.
S. Li, Y. Yao, J. Hu, G. Liu, X. Yao, and J. Hu, “An Ensemble Stacked Convolutional Neural Network Model for Environmental Event Sound Recognition,” Appl. Sci. 2018, 8, 1152, doi: 10.3390/app8071152
2018
Later among the works it cites.
Q. Kong, Y. Xu, W. Wang, MD. Plumbley, “Sound Event Detection and Time-Frequency Segmentation from Weakly Labelled Data,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP) , vol. 27, no. 4, pp. 777-787, 2019. Page 777-787
2019
Closest in time.