Fetching the paper…
Reading the bibliography…
Various neural network-based approaches have been proposed for more robust and accurate voice activity detection (VAD).
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, and V. Zue, “Timit acoustic-phonetic continuous speech corpus,” Linguistic Data Consortium , 11 1992
1992
Earlier work this paper cites.
J. Sohn, N. S. Kim, and W. Sung, “A statistical model-based voice activity detection,” IEEE Signal Processing Letters , vol. 6, no. 1, 1999
1999
Earlier work this paper cites.
X.-L. Zhang and J. Wu, “Deep belief networks based voice activity detection,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 21, no. 4, 2013
2013
Earlier work this paper cites.
T. Hughes and K. Mierle, “Recurrent neural networks for voice activity detection,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing , 2013
2013
Earlier work this paper cites.
X.-L. Zhang and D. Wang, “Boosting contextual information for deep neural network based voice activity detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 2, 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision , 2016
2016
Earlier work this paper cites.
X.-L. Zhang and D. Wang, “Boosting contextual information for deep neural network based voice activity detection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 2, 2016
2016
Earlier work this paper cites.
E. Real, S. Moore, A. Selle, S. Saxena, Y. L. Suematsu, J. Tan, Q. V. Le, and A. Kurakin, “Large-scale evolution of image classifiers,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 70. PMLR, 2017
2017
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 70. PMLR, 2017
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing , 2017
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “SGDR: Stochastic Gradient Descent with Warm Restarts,” in International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
J. Kim and M. Hahn, “Voice activity detection using an adaptive context attention model,” IEEE Signal Processing Letters , vol. 25, no. 8, 2018
2018
Cited alongside, same era.
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le, “Learning transferable architectures for scalable image recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Cited alongside, same era.
H. Pham, M. Guan, B. Zoph, Q. Le, and J. Dean, “Efficient neural architecture search via parameters sharing,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80. PMLR, 2018
2018
Cited alongside, same era.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollar, “Designing network design spaces,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2020
2020
Later among the works it cites.
J. Li, C. Liang, B. Zhang, Z. Wang, F. Xiang, and X. Chu, “Neural Architecture Search on Acoustic Scene Classification,” in Proc. Interspeech 2020 , 2020
2020
Later among the works it cites.
Y.-C. Chen, J.-Y. Hsu, C.-K. Lee, and H. yi Lee, “DARTS-ASR: Differentiable Architecture Search for Multilingual Speech Recognition and Adaptation,” in Proc. Interspeech 2020 , 2020
2020
Later among the works it cites.
J. Kim, J. Wang, S. Kim, and Y. Lee, “Evolved Speech-Transformer: Applying Neural Architecture Search to End-to-End Automatic Speech Recognition,” in Proc. Interspeech 2020 , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Cited alongside, same era.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Cited alongside, same era.
S. Chaudhuri, J. Roth, D. P. W. Ellis, A. Gallagher, L. Kaver, R. Marvin, C. Pantofaru, N. Reale, L. Guarino Reid, K. Wilson, and Z. Xi, “AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies,” in Proc. Interspeech 2018 , 2018
2018
Cited alongside, same era.
T. Elsken, J. H. Metzen, and F. Hutter, “Neural architecture search: A survey,” Journal of Machine Learning Research , vol. 20, no. 55, pp. 1–21, 2019. [Online]. Available: http://jmlr.org/papers/v20/18-598.html
2019
Cited alongside, same era.
H. Zhou, M. Yang, J. Wang, and W. Pan, “BayesNAS: A Bayesian approach for neural architecture search,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019
2019
Cited alongside, same era.
H. Liu, K. Simonyan, and Y. Yang, “DARTS: Differentiable architecture search,” in International Conference on Learning Representations , 2019
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech 2019 , 2019
2019
Cited alongside, same era.
Z.-H. Tan, A. kr. Sarkar, and N. Dehak, “rVAD: An unsupervised segment-based robust voice activity detection method,” Computer Speech & Language , vol. 59, 2020
2020
Cited alongside, same era.
L. Li and A. Talwalkar, “Random search and reproducibility for neural architecture search,” in Proceedings of The 35th Uncertainty in Artificial Intelligence Conference , ser. Proceedings of Machine Learning Research, vol. 115. PMLR, 22–25 Jul 2020
2020
Later among the works it cites.
S. Ding, T. Chen, X. Gong, W. Zha, and Z. Wang, “AutoSpeech: Neural Architecture Search for Speaker Recognition,” in Proc. Interspeech 2020 , 2020
2020
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020) , 2020
2020
Later among the works it cites.
Y. Chen, H. Dinkel, M. Wu, and K. Yu, “Voice Activity Detection in the Wild via Weakly Supervised Sound Event Detection,” in Proc. Interspeech 2020 , 2020
2020
Later among the works it cites.
Y. R. Jo, Y. Ki Moon, W. I. Cho, and G. Sik Jo, “Self-attentive vad: Context-aware detection of voice from noise,” in 2021 IEEE International Conference on Acoustics, Speech and Signal Processing , 2021
2021
Later among the works it cites.
J. Mellor, J. Turner, A. Storkey, and E. J. Crowley, “Neural architecture search without training,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 139. PMLR, 2021
2021
Later among the works it cites.
B. X. Ru, X. Wan, X. Dong, and M. A. Osborne, “Interpretable neural architecture search via bayesian optimisation with weisfeiler-lehman kernels,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
C. K. Reddy, H. Dubey, K. Koishida, A. Nair, V. Gopal, R. Cutler, S. Braun, H. Gamper, R. Aichner, and S. Srinivasan, “INTERSPEECH 2021 Deep Noise Suppression Challenge,” in Proc. Interspeech 2021 , 2021
2021
Later among the works it cites.