Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) achieves great success in speech recognition, while limited exploration has been attempted for other speech processing tasks.
J. Allen and D. Berkley, “Image method for efficiently simulating small-room acoustics,” The Journal of the Acoustical Society of America (JASA) , vol. 65, pp. 943–950, 1979
1979
Earlier work this paper cites.
L. D. C. Philadelphia, “CSR-II (WSJ1) Complete,” 1994, http://catalog.ldc.upenn.edu/LDC94S13A
1994
Earlier work this paper cites.
M. Przybocki and A. Martin, “2000 nist speaker recognition evaluation (ldc2001s97),” in Philadelphia, New Jersey: Linguistic Data Consortium , 2001
2001
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in International Conference on Machine Learning (ICML) , 2006, pp. 369–376
2006
Earlier work this paper cites.
E. A. Habets and S. Gannot, “Generating sensor signals in isotropic noise fields,” The Journal of the Acoustical Society of America (JASA) , vol. 122, no. 6, pp. 3464–3470, 2007
2007
Earlier work this paper cites.
Z. N. Karam, W. M. Campbell, and N. Dehak, “Towards reduced false-alarms using cohorts,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2011, pp. 4512–4515
2011
Earlier work this paper cites.
S. Cumani, P. D. Batzu, D. Colibro, C. Vair, P. Laface, and V. Vasilakakis, “Comparison of speaker recognition approaches for real applications.” in Interspeech , 2011, pp. 2365–2368
2011
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 22, no. 12, pp. 1849–1858, 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep networks with stochastic depth,” in European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 646–661
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Learning latent representations for speech generation and transformation,” in Interspeech , 2017, pp. 1273–1277
2017
Earlier work this paper cites.
——, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 5220–5224
2017
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Deep recurrent networks for separation and recognition of single-channel speech in nonstationary background audio,” in New Era for Robust Speech Recognition . Springer, 2017, pp. 165–186
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Narayanan, A. Misra, K. C. Sim, G. Pundak, A. Tripathi, M. Elfeky, P. Haghani, T. Strohman, and M. Bacchiani, “Toward domain-invariant speech recognition via large scale training,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 441–447
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Yoshioka, H. Erdogan, Z. Chen, and F. Alleva, “Multi-microphone neural speech separation for far-field multi-talker speech recognition,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5739–5743
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations (ICLR) , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics (NAACL) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in Advances in Neural Information Processing Systems (NeurIPS) , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019
2019
Earlier work this paper cites.
A. Narayanan, R. Prabhavalkar, C.-C. Chiu, D. Rybach, T. N. Sainath, and T. Strohman, “Recognizing long-form speech using streaming end-to-end models,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 920–927
2019
Earlier work this paper cites.
Y.-C. Chen, S.-F. Huang, H.-y. Lee, Y.-H. Wang, and C.-H. Shen, “Audio word2vec: Sequence-to-sequence autoencoding for unsupervised learning of audio segmentation and representation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 27, no. 9, pp. 1481–1493, 2019
2019
Earlier work this paper cites.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP) , vol. 27, no. 12, pp. 2041–2053, 2019
2019
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Interspeech , 2019, pp. 146–150
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition.” in Interspeech , 2019
2019
Earlier work this paper cites.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,” in Interspeech , 2019, pp. 161–165
2019
Earlier work this paper cites.
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 4690–4699
2019
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with permutation-free objectives,” in Interspeech , 2019, pp. 4300–4304
2019
Cited alongside, same era.
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” in Interspeech , 2019, pp. 2613–2617
2019
Cited alongside, same era.
2021
Closest in time.
C. Wang, Y. Wu, Y. Qian, K. Kumatani, S. Liu, F. Wei, M. Zeng, and X. Huang, “Unispeech: Unified speech representation learning with labeled and unlabeled data,” in International Conference on Machine Learning(ICML) , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 2021, pp. 10 937–10 947. [Online]. Available: http://proceedings.mlr.press/v139/wang21y.html
2021
Closest in time.
S. Yang, P. Chi, Y. Chuang, C. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G. Lin, T. Huang, W. Tseng, K. Lee, D. Liu, Z. Huang, S. Dong, S. Li, S. Watanabe, A. Mohamed, and H. Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Interspeech , 2021, pp. 1194–1198
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
V. Pratap et al. , “Wav2letter++: A fast open-source speech recognition system,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6460–6464
2019
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research (JMLR) , vol. 21, no. 140, pp. 1–67, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Cited alongside, same era.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7669–7673
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, and J. Li, “Continuous speech separation: Dataset and analysis,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7284–7288
2020
Cited alongside, same era.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
G. Chen, S. Chai, G. Wang, J. Du, W. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Y. Wang, Z. You, and Z. Yan, “Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,” in Interspeech , 2021
2021
Closest in time.
2021
Closest in time.
S. Chen, Y. Wu, Z. Chen, J. Wu, J. Li, T. Yoshioka, C. Wang, S. Liu, and M. Zhou, “Continuous speech separation with conformer,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5749–5753
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Wu, Z. Chen, S. Chen, Y. Wu, T. Yoshioka, N. Kanda, S. Liu, and J. Li, “Investigation of Practical Aspects of Single Channel Speech Separation for ASR,” in Interspeech , 2021, pp. 3066–3070
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Thienpondt, B. Desplanques, and K. Demuynck, “The idlab voxsrc-20 submission: Large margin fine-tuning and quality-aware score calibration in dnn based speaker verification,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5814–5818
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
S. Chen, Y. Wu, Z. Chen, T. Yoshioka, S. Liu, J. Li, and X. Yu, “Don’t shoot butterfly with rifles: Multi-channel continuous speech separation with early exit transformer,” in International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6139–6143
2021
Closest in time.
S. Chen, Y. Wu, Z. Chen, J. Wu, T. Yoshioka, S. Liu, J. Li, and X. Yu, “Ultra Fast Speech Separation Model with Teacher Student Learning,” in Interspeech , 2021, pp. 3026–3030
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2022
Closest in time.
J. Bai, B. Li, Y. Zhang, A. Bapna, N. Siddhartha, K. C. Sim, and T. N. Sainath, “Joint unsupervised and supervised training for multilingual asr,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6402–6406
2022
Closest in time.
F. Landini, J. Profant, M. Diez, and L. Burget, “Bayesian hmm clustering of x-vector sequences (vbx) in speaker diarization: theory, implementation and analysis on standard tasks,” Computer Speech & Language , vol. 71, p. 101254, 2022
2022
Closest in time.