Fetching the paper…
Reading the bibliography…
Music, speech, and acoustic scene sound are often handled separately in the audio domain because of their different signal characteristics.
Evaluation of algorithms using games: The case of music tagging
Edith Law, Kris West, Michael I Mandel, Mert Bay, and J Stephen Downie · 2009
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Earlier work this paper cites.
End-to-end learning for music audio
Sander Dieleman and Benjamin Schrauwen · 2014
Earlier work this paper cites.
Learning the speech front-end with raw waveform cldnns
Tara N Sainath, Ron J Weiss, Andrew Senior, Kevin W Wilson, and Oriol Vinyals · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Experimenting with musically motivated convolutional neural networks
Jordi Pons, Thomas Lidy, and Xavier Serra · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sound event detection using weakly labeled dataset with stacked convolutional and recurrent neural network
Sharath Adavanne and Tuomas Virtanen · 2017
Cited alongside, same era.
Stacked convolutional and recurrent neural networks for bird audio detection
Sharath Adavanne, Konstantinos Drossos, Emre Çakır, and Tuomas Virtanen · 2017
Cited alongside, same era.
Stacked convolutional and recurrent neural networks for music emotion recognition
Miroslav Malik, Sharath Adavanne, Konstantinos Drossos, Tuomas Virtanen, Dasa Ticha, and Roman Jarina · 2017
Cited alongside, same era.
Very deep convolutional neural networks for raw waveforms
Wei Dai, Chia Dai, Shuhui Qu, Juncheng Li, and Samarjit Das · 2017
Cited alongside, same era.
Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam · 2017
Cited alongside, same era.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2017
Closest in time.
Multi-level and multi-scale feature aggregation using pretrained convolutional neural networks for music auto-tagging
Jongpil Lee and Juhan Nam · 2017
Closest in time.
Speech commands: A public dataset for single-word speech recognition
Pete Warden · 2017
Closest in time.
Dcase 2017 challenge setup: tasks, datasets and baseline system
Annamaria Mesaros, Toni Heittola, Aleksandr Diment, Benjamin Elizalde, Ankit Shah, Emmanuel Vincent, Bhiksha Raj, and Tuomas Virtanen · 2017
Closest in time.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combining multi-scale features using sample-level deep convolutional neural networks for weakly supervised sound event detection
Jongpil Lee, Jiyoung Park, Sangeun Kum, Youngho Jeong, and Juhan Nam · 2017
Cited alongside, same era.
Sample-level cnn architectures for music auto-tagging using raw waveforms
Taejun Kim, Jongpil Lee, and Juhan Nam · 2017
Cited alongside, same era.
https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/leaderboard
TensorFlow Speech Recognition Challenge
Cited in the paper.
Yong Xu, Qiuqiang Kong, Wenwu Wang, and Mark D Plumbley · 2017
Closest in time.