Fetching the paper…
Reading the bibliography…
End-to-end learning models using raw waveforms as input have shown superior performances in many audio recognition tasks.
“Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position,”
Kunihiko Fukushima, · 1980
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Learning the speech front-end with raw waveform cldnns.,”
Tara N Sainath, Ron J Weiss, Andrew W Senior, Kevin W Wilson, and Oriol Vinyals, · 2015
Earlier work this paper cites.
“Identity mappings in deep residual networks,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Sample-level deep convolutional neural networks for music auto-tagging using raw waveforms,”
Jongpil Lee, Jiyoung Park, Keunhyoung Luke Kim, and Juhan Nam, · 2017
Earlier work this paper cites.
“Very deep convolutional neural networks for raw waveforms,”
Wei Dai, Chia Dai, Shuhui Qu, Juncheng Li, and Samarjit Das, · 2017
Earlier work this paper cites.
Human and Machine Hearing: Extracting Meaning from Sound
Richard F. Lyon, · 2017
Cited alongside, same era.
“Automatic differentiation in PyTorch,”
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer, · 2017
Cited alongside, same era.
“End-to-end speech emotion recognition using deep neural networks,”
Panagiotis Tzirakis, Jiehao Zhang, and Bjorn W. Schuller, · 2018
Cited alongside, same era.
“Squeeze-and-excitation networks,”
Jie Hu, Li Shen, and Gang Sun, · 2018
Cited alongside, same era.
“Speech commands: A dataset for limited-vocabulary speech recognition,”
Pete Warden, · 2018
Cited alongside, same era.
“SampleCNN: End-to-end deep convolutional neural networks using very small filters for music classification,”
“A neural attention model for speech command recognition,”
Douglas Coimbra de Andrade, Sabato Leo, Martin Loesener Da Silva Viana, and Christoph Bernkopf, · 2018
Later among the works it cites.
“Comparison and analysis of SampleCNN architectures for audio classification,”
Taejun Kim, Jongpil Lee, and Juhan Nam, · 2019
Closest in time.
“Direct modelling of speech emotion from raw speech,”
Siddique Latif, Rajib Rana, Sara Khalifa, Raja Jurdak, and Julien Epps, · 2019
Closest in time.
“End-to-end environmental sound classification using a 1d convolutional neural network,”
Sajjad Abdoli, Patrick Cardinal, and Alessandro Lameiras Koerich, · 2019
Closest in time.
“Data-driven harmonic filters for audio representation learning,”
Minz Won, Sanghyuk Chun, Oriol Nieto, and Xavier Serrc, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jongpil Lee, Jiyoung Park, Keunhyoung Kim, and Juhan Nam, · 2018
Cited alongside, same era.