Fetching the paper…
Reading the bibliography…
Automatic Speech Recognition (ASR) systems generalize poorly on accented speech.
J. J. Humphries, P. C. Woodland, and D. Pearce, “Using accent-specific pronunciation modelling for robust speech recognition,” in Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP’96 , vol. 4. IEEE, 1996, pp. 2324–2327
1996
Earlier work this paper cites.
J. J. Humphries and P. C. Woodland, “The use of accent-specific pronunciation dictionaries in acoustic model training,” in Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP’98 (Cat. No. 98CH36181) , vol. 1. IEEE, 1998, pp. 317–320
1998
Earlier work this paper cites.
C. Huang, T. Chen, S. Li, E. Chang, and J. Zhou, “Analysis of speaker variability,” in Seventh European Conference on Speech Communication and Technology , 2001
2001
Earlier work this paper cites.
S. Goronzy, S. Rapp, and R. Kompe, “Generating non-native pronunciation variants for lexicon adaptation,” Speech Communication , vol. 42, no. 1, pp. 109–123, 2004
2004
Earlier work this paper cites.
C. Huang, T. Chen, and E. Chang, “Accent issues in large vocabulary continuous speech recognition,” International Journal of Speech Technology , vol. 7, no. 2, pp. 141–153, 2004
2004
Earlier work this paper cites.
D. Vergyri, L. Lamel, and J.-L. Gauvain, “Automatic speech recognition of multiple accented english data,” in Eleventh Annual Conference of the International Speech Communication Association , 2010
2010
Earlier work this paper cites.
L. Loots and T. Niesler, “Automatic conversion between pronunciations of different english accents,” Speech Communication , vol. 53, no. 1, pp. 75–84, 2011
2011
Earlier work this paper cites.
J. Holmes, An introduction to sociolinguistics . Routledge, 2013
2013
Earlier work this paper cites.
M. Najafian, A. DeMarco, S. Cox, and M. Russell, “Unsupervised model selection for recognition of regional accented speech,” in Fifteenth annual conference of the international speech communication association , 2014
2014
Earlier work this paper cites.
Y. Huang, D. Yu, C. Liu, and Y. Gong, “Multi-accent deep neural network acoustic model with accent-specific top layer using the kld-regularized model adaptation,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Earlier work this paper cites.
M. Lehr, K. Gorman, and I. Shafran, “Discriminative pronunciation modeling for dialectal speech recognition,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Earlier work this paper cites.
M. Chen, Z. Yang, J. Liang, Y. Li, and W. Liu, “Improving deep neural networks based multi-accent mandarin speech recognition using i-vectors and accent-specific top layer,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “Achieving human parity in conversational speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. PP, 10 2016
2016
Earlier work this paper cites.
Y. Shinohara, “Adversarial multi-task learning of deep neural networks for robust speech recognition,” in Interspeech 2016 , 2016, pp. 2369–2372. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2016-879
2016
Earlier work this paper cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,” J. Mach. Learn. Res. , vol. 17, no. 1, p. 2096–2030, Jan. 2016
2016
Cited alongside, same era.
K. Rao and H. Sak, “Multi-accent speech recognition with hierarchical grapheme based models,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2017, pp. 4815–4819
2017
Cited alongside, same era.
X. Yang, K. Audhkhasi, A. Rosenberg, S. Thomas, B. Ramabhadran, and M. Hasegawa-Johnson, “Joint modeling of accents and acoustics for multi-accent speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 1–5
2018
Cited alongside, same era.
T. Fukuda, R. Fernandez, A. Rosenberg, S. Thomas, B. Ramabhadran, A. Sorin, and G. Kurata, “Data augmentation improves recognition of foreign accented speech.” in Interspeech , 2018, pp. 2409–2413
S. Yoo, I. Song, and Y. Bengio, “A highly adaptive acoustic model for accurate multi-dialect speech recognition,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 5716–5720
2019
Later among the works it cites.
A. Jain, V. P. Singh, and S. P. Rath, “A multi-accent acoustic model using mixture of experts for speech recognition.” in INTERSPEECH , 2019, pp. 779–783
2019
Later among the works it cites.
T. Viglino, P. Motlicek, and M. Cernak, “End-to-end accented speech recognition.” in INTERSPEECH , 2019, pp. 2140–2144
2019
Later among the works it cites.
J. Shor, D. Emanuel, O. Lang, O. Tuval, M. Brenner, J. Cattiau, F. Vieira, M. McNally, T. Charbonneau, M. Nollstadt, A. Hassidim, and Y. Matias, “Personalizing ASR for Dysarthric and Accented Speech with Limited Data,” in Proc. Interspeech 2019 , 2019, pp. 784–788. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-1427
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
A. Jain, M. Upreti, and P. Jyothi, “Improved accented speech recognition using accent embeddings and multi-task learning.” in Interspeech , 2018, pp. 2454–2458
2018
Cited alongside, same era.
B. Li, T. N. Sainath, K. C. Sim, M. Bacchiani, E. Weinstein, P. Nguyen, Z. Chen, Y. Wu, and K. Rao, “Multi-dialect speech recognition with a single sequence-to-sequence model,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 4749–4753
2018
Cited alongside, same era.
S. Ghorbani and J. H. Hansen, “Leveraging native language information for improved accented speech recognition,” in Proc. Interspeech 2018 , 2018, pp. 2449–2453. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1378
2018
Cited alongside, same era.
J. Schwarz, W. Czarnecki, J. Luketina, A. Grabska-Barwinska, Y. W. Teh, R. Pascanu, and R. Hadsell, “Progress & compress: A scalable framework for continual learning,” in International Conference on Machine Learning . PMLR, 2018, pp. 4528–4537
2018
Cited alongside, same era.
M. Grace, M. Bastani, and E. Weinstein, “Occam’s adaptation: A comparison of interpolation of bases adaptation methods for multi-dialect acoustic modeling with lstms,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 174–181
2018
Cited alongside, same era.
S. Sun, C.-F. Yeh, M. Hwang, M. Ostendorf, and L. Xie, “Domain adversarial training for accented speech recognition,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 4854–4858, 2018
2018
Cited alongside, same era.
S. Ghorbani, S. Khorram, and J. H. Hansen, “Domain expansion in dnn-based acoustic models for robust speech recognition,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2019, pp. 107–113
2019
Cited alongside, same era.
G. I. Winata, Z. Lin, and P. Fung, “Learning multilingual meta-embeddings for code-switching named entity recognition,” in Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019) , 2019, pp. 181–186
2019
Cited alongside, same era.
A. Koenecke, A. Nam, E. Lake, J. Nudell, M. Quartey, Z. Mengesha, C. Toups, J. R. Rickford, D. Jurafsky, and S. Goel, “Racial disparities in automated speech recognition,” Proceedings of the National Academy of Sciences , vol. 117, no. 14, pp. 7684–7689, 2020
2020
Later among the works it cites.
A. Prasad and P. Jyothi, “How accents confound: Probing for accent information in end-to-end speech recognition systems,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 3739–3753
2020
Later among the works it cites.
B. Houston and K. Kirchhoff, “Continual learning for multi-dialect acoustic models,” Proc. Interspeech 2020 , pp. 576–580, 2020
2020
Later among the works it cites.
M. T. Turan, E. Vincent, and D. Jouvet, “Achieving multi-accent asr via unsupervised acoustic model adaptation,” in INTERSPEECH 2020 , 2020
2020
Later among the works it cites.
W. Rao, J. Zhang, and J. Wu, “Improved blstm rnn based accent speech recognition using multi-task learning and accent embeddings,” in Proceedings of the 2020 2nd International Conference on Image, Video and Signal Processing , 2020, pp. 1–6
2020
Later among the works it cites.
Y. Chen, Z. Yang, C. feng Yeh, M. Jain, and M. L. Seltzer, “Aipnet: Generative adversarial pre-training of accent-invariant networks for end-to-end speech recognition,” ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6979–6983, 2020
2020
Later among the works it cites.
K. Khandelwal, P. Jyothi, A. Awasthi, and S. Sarawagi, “Black-Box Adaptation of ASR for Accented Speech,” in Proc. Interspeech 2020 , 2020, pp. 1281–1285. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-3162
2020
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020) , 2020, pp. 4211–4215
2020
Later among the works it cites.
J. Meyer, L. Rauchenstein, J. D. Eisenberg, and N. Howell, “Artie bias corpus: An open dataset for detecting demographic bias in speech applications,” in Proceedings of The 12th Language Resources and Evaluation Conference , 2020, pp. 6462–6468
2020
Later among the works it cites.