Fetching the paper…
Reading the bibliography…
We propose a new deep network for audio event recognition, called AENet.
W. Fisher, G. Doddington, and K. Goudie-Marshall, “The DARPA speech recognition research database: Specifications and status,” in Proc. DARPA Workshop on Speech Recognition , 1986, pp. 93–99
1986
Earlier work this paper cites.
H. David, “A tractable inference algorithm for diagnosing multiple diseases.” in Proc. UAI , 1989, pp. 163–171
1989
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient based learning applied to document recognition,” in Proc. of the IEEE , vol. 86, no. 11, 1998, pp. 2278–2324
1998
Earlier work this paper cites.
S. Nakamura, K. Hiyane, and F. Asano, “Acoustical sound database in real environments for sound scene understanding and hands-free speech recognition,” in International Conference on Language Resources & Evaluation , 2000, pp. 2–5
2000
Earlier work this paper cites.
T. Zhang and C. C. Jay Kuo, “Audio content analysis for online audiovisual data segmentation and classification,” IEEE Transactions on Speech and Audio Processing , vol. 9, no. 4, pp. 441–457, 2001
2001
Earlier work this paper cites.
M. R. Naphade and T. S. Huang, “Indexing , Filtering , and Retrieval,” IEEE Transactions on Multimedia , vol. 3, no. 1, pp. 141–151, 2001
2001
Earlier work this paper cites.
P. Viola, J. C. Platt, and C. Zhang, “Multiple instance boosting for object detection,” in Proc. NIPS , vol. 74, no. 10, 2005, pp. 1769–1775
2005
Earlier work this paper cites.
A. J. Eronen, V. T. Peltonen, J. T. Tuomi, A. P. Klapuri, S. Fagerlund, T. Sorsa, G. Lorho, and J. Huopaniemi, “Audio-based context recognition,” IEEE Transaction on Audio, Speech and Language Processing , vol. 14, no. 1, pp. 321–329, 2006
2006
Earlier work this paper cites.
A. Temko, E. Monte, and C. Nadeu, “Comparison of sequence discriminant support vector machines for acoustic event classification,” in Proc. ICASSP , vol. 5, 2006, pp. 721–724
2006
Earlier work this paper cites.
W. Choi, S. Park, D. K. Han, and H. Ko, “Acoustic event recognition using dominant spectral basis vectors,” in Proc. Interspeech , 2015, pp. 2002–2006
2006
Earlier work this paper cites.
D. Mostefa, N. Moreau, K. Choukri, G. Potamianos, S. M. Chu, A. Tyagi, J. R. Casas, J. Turmo, L. Cristoforetti, F. Tobia, A. Pnevmatikakis, V. Mylonakis, F. Talantzis, S. Burger, R. Stiefelhagen, K. Bernardin, and C. Rochet, “The CHIL audiovisual corpus for lecture and meeting analysis inside smart rooms,” Language Resources and Evaluation , vol. 41, no. 3-4, p. 389–407, 2007
2007
Earlier work this paper cites.
M. Xu, C. Xu, L. Duan, J. S. Jin, and S. Luo, “Audio keywords generation for sports video analysis,” ACM Transaction on Multimedia Computing, Communications, and Applications , vol. 4, no. 2, pp. 1–23, 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR , 2009
2009
Earlier work this paper cites.
S. Chu, S. Narayanan, and C.-C. J.Kuo, “Environmental sound recognition with time frequency audio features,” IEEE Transaction on Audio, Speech and Language Processing , vol. 17, no. 6, pp. 1142–1158, 2009
2009
Earlier work this paper cites.
K. Lee and D. P. W. Ellis, “Audio-based semantic concept classification for consumer video,” IEEE Transactions on Audio, Speech and Language Processing , vol. 18, no. 6, pp. 1406–1416, 2010
2010
Earlier work this paper cites.
X. Zhuang, X. Zhou, M. a. Hasegawa-Johnson, and T. S. Huang, “Real-world acoustic event detection,” Pattern Recognition Letters , vol. 31, no. 12, pp. 1543–1551, 2010
2010
Earlier work this paper cites.
W. Hu, N. Xie, L. Li, X. Zeng, and S. Maybank, “A Survey on Visual Content-Based Video Indexing and Retrieval,” IEEE Transactions on Systems, Man, and Cybernetics, Part C , vol. 41, no. 6, pp. 797–819, 2011
2011
Earlier work this paper cites.
NIST, “2011 trecvid multimedia event detection evalua- tion plan,” http://www.nist.gov/itl/iad/mig/ upload/MED11-EvalPlan-V03-20110801a.pdf
2011
Earlier work this paper cites.
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu, “Action recognition by dense trajectories,” in Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on . IEEE, 2011, pp. 3169–3176
2011
Earlier work this paper cites.
S. Pancoast and M. Akbacak, “Bag-of-Audio-Words Approach for Multimedia Event Classification,” in Proc. Interspeech , Portland, OR, USA, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Proc. NIPS , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition,” Signal Processing Magazine , 2012
2012
Earlier work this paper cites.
O. Abdel-Hamid, A. Mohamed, H. Jiang, and G. Penn, “Applying convolutional neural networks concepts to hybrid NN-HMM model for Speech Recognition,” in ICASSP , 2012, pp. 4277–4280
2012
Cited alongside, same era.
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, “Improving neural networks by preventing co-adaptation of feature detectors,” arXiv: 1207.0580 , 2012
2012
Cited alongside, same era.
Q. Wu, Z. Wang, F. Deng, Z. Chi, and D. D. Feng, “Realistic human action recognition with multimodal feature selection and fusion,” IEEE Transactions on Systems, Man, and Cybernetics Part A:Systems and Humans , vol. 43, no. 4, pp. 875–885, 2013
2013
Cited alongside, same era.
D. Oneata, J. Verbeek, and C. Schmid, “Action and event recognition with fisher vectors on a compact feature set,” in Proc. ICCV , December 2013
2013
Cited alongside, same era.
O. Gencoglu, T. Virtanen, and H. Huttunen, “Recognition of acoustic events using deep neural networks,” Proc. EUSIPCO , no. 1-5 Sept. 2014, pp. 506–510, 2014
2014
Later among the works it cites.
J. Liang, Q. Jin, X. He, Y. Gang, J. Xu, and X. Li, “Detecting semantic concepts in consumer videos using audio,” in ICASSP , 2015, pp. 2279–2283
2015
Later among the works it cites.
M. Gygli, H. Grabner, and L. Van Gool, “Video summarization by learning submodular mixtures of objectives,” in CVPR , 2015
2015
Later among the works it cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proc. ICLR , 2015, pp. 1–14
2015
Later among the works it cites.
J. Y.-H. Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici, “Beyond Short Snippets: Deep Networks for Video Classification,” in Proc. CVPR , 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. (2013) Decaf: A deep convolutional activation feature for generic visual recognition
2013
Cited alongside, same era.
P. Over, J. Fiscus, and G. Sanders, “Trecvid 2013 – an introduction to the goals, tasks, data, eval- uation mechanisms, and metrics,” http://www- nlpir.nist.gov/projects/tv2013/
2013
Cited alongside, same era.
H. Wang and C. Schmid, “Action recognition with improved trajectories,” in Proceedings of the IEEE International Conference on Computer Vision , 2013, pp. 3551–3558
2013
Cited alongside, same era.
Z. Huang, Y. C. Cheng, K. Li, V. Hautamäki, and C. H. Lee, “A blind segmentation approach to acoustic event detection based on I-vector,” in Proc. Interspeech , 2013, pp. 2282–2286
2013
Cited alongside, same era.
O. Abdel-hamid, L. Deng, and D. Yu, “Exploring Convolutional Neural Network Structures and Optimization Techniques for Speech Recognition,” in Interspeech , 2013, pp. 3366–3370
2013
Cited alongside, same era.
N. Jaitly and G. E. Hinton, “Vocal tract length perturbation (VTLP) improves speech recognition,” ICML , vol. 28, 2013
2013
Cited alongside, same era.
F. Font, G. Roma, and X. Serra, “Freesound technical demo,” in Proc. 21st ACM international conference on Multimedia , 2013
2013
Cited alongside, same era.
M. Sun, A. Farhadi, and S. Seitz, “Ranking domain-specific highlights by analyzing edited videos,” in ECCV , 2014
2014
Cited alongside, same era.
2015
Later among the works it cites.
L. Wang, Y. Qiao, and X. Tang, “Action recognition with trajectory-pooled deep-convolutional descriptors,” in CVPR , 2015
2015
Later among the works it cites.
M. Cimpoi, S. Maji, and A. Vedaldi, “Deep filter banks for texture recognition and segmentation,” in Proc. CVPR , 2015
2015
Later among the works it cites.
H. Phan, L. Hertel, M. Maass, R. Mazur, and A. Mertins, “Representing nonspeech audio signals through speech classification models,” in Proc. Interspeech , 2015, pp. 3441–3445
2015
Later among the works it cites.
J. Beltrán, E. Chávez, and J. Favela, “Scalable identification of mixed environmental sounds, recorded from heterogeneous sources,” Pattern Recognition Letters , vol. 68, pp. 153–160, 2015
2015
Later among the works it cites.
X. Lu, P. Shen, Y. Tsao, C. Hori, and H. Kawai, “Sparse representation with temporal max-smoothing for acoustic event detection,” in Proc. Interspeech , 2015, pp. 1176–1180
2015
Later among the works it cites.
H. Lim, M. J. Kim, and H. Kim, “Robust sound event classification using LBP-HOG based bag-of-audio-words feature representation,” in Proc. Interspeech , 2015, pp. 3325–3329
2015
Later among the works it cites.
K. Ashraf, B. Elizalde, F. Iandola, M. Moskewicz, J. Bernd, G. Friedland, and K. Keutzer, “Audio-based multimedia event detection with DNNs and Sparse Sampling Categories and Subject Descriptors,” in Proc. ICMR , 2015, pp. 611–614
2015
Later among the works it cites.
M. Espi, M. Fujimoto, K. Kinoshita, and T. Nakatani, “Exploiting spectro-temporal locality in deep learning based acoustic event detection,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 26, pp. 1–12, 2015
2015
Later among the works it cites.
J. Wu, Yinan Yu, Chang Huang, and Kai Yu, “Deep multiple instance learning for image classification and auto-annotation,” in Proc. CVPR , 2015, pp. 3460–3469
2015
Later among the works it cites.
E. Battenberg, S. Dieleman, D. Nouri, E. Olson, C. Raffel, J. Schlüter, S. K. Sønderby, D. Maturana, M. Thoma, and et al., “Lasagne: First release.” http://dx.doi.org/10.5281/zenodo.27878, Aug. 2015
2015
Later among the works it cites.
N. Takahashi, M. Gygli, B. Pfister, and L. Van Gool, “Deep Convolutional Neural Networks and Data Augmentation for Acoustic Event Detection,” in Proc. Interspeech , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
M. Gygli, Y. Song, and L. Cao, “Video2gif: Automatic generation of animated gifs from video,” in CVPR , 2016
2016
Later among the works it cites.
F. M. Yun Wang, Leonardo Neves, “Audio-based multimedia event detection using deep recurent neural nework,” in ICASSP , 2016, pp. 2742–2746
2016
Later among the works it cites.
T. Sercu, C. Puhrsch, B. Kingsbury, and Y. LeCun, “Very deep multilingual convolutional neural networks for LVCSR,” in Proc. ICASSP , 2016, pp. 4955–4959
2016
Later among the works it cites.
K. Soomro, A. R. Zamir, and M. Shah, “UCF101: A Dataset of 101 Human Action Classes From Videos in The Wild,” in CRCV-TR-12-01 , 2012
2016
Later among the works it cites.