Fetching the paper…
Reading the bibliography…
In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization of non-standard words) is imputed by separate post-processing models.
K. H. Albrow, “The english writing system: Notes towards a description,” Schools Council Program in Linguistics and English Teaching, papers series 2 , no. 2, 1972
1972
Earlier work this paper cites.
K. Fukushima, “Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position,” Biological Cybernetics , vol. 36, pp. 193–202, 1980
1980
Earlier work this paper cites.
D. B. Paul and J. M. Baker, “The design for the Wall Street Journal-based CSR corpus,” in Speech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23-26, 1992 , 1992
1992
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in ICASSP , 1992
1992
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren, “DARPA TIMIT acoustic phonetic continuous speech corpus,” 1993
1993
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The Fisher Corpus: a resource for the next generations of speech-to-text,” in Proceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04) . Lisbon, Portugal: European Language Resources Association (ELRA), May 2004. [Online]. Available: http://www.lrec-conf.org/proceedings/lrec2004/pdf/767.pdf
2004
Earlier work this paper cites.
A. Graves, S. Fernández, and F. Gomez, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in In Proceedings of the International Conference on Machine Learning, ICML 2006 , 2006, pp. 369–376
2006
Earlier work this paper cites.
S. Peitz, M. Freitag, A. Mauser, and H. Ney, “Modeling punctuation prediction as machine translation,” in in Proceedings of the International Workshop on Spoken Language Translation (IWSLT , 2011
2011
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . Brussels, Belgium: Association for Computational Linguistics, Nov. 2018, pp. 66–71. [Online]. Available: https://www.aclweb.org/anthology/D18-2012
2012
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in ICML 29 , 2012
2012
Earlier work this paper cites.
J. Le Roux and E. Vincent, “A categorization of robust speech processing datasets,” Mitsubishi Electric Research Labs TR2014-116, Technical Report, 2014
2014
Earlier work this paper cites.
O. Abdel-Hamid, A. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 22, no. 10, pp. 1533–1545, 2014. [Online]. Available: https://doi.org/10.1109/TASLP.2014.2339736
2014
Earlier work this paper cites.
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V. Chaudhary, M. Young, J.-F. Crespo, and D. Dennison, “Hidden technical debt in machine learning systems,” in Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015. [Online]. Available: https://proceedings.neurips.cc/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
M. S. C. International, “The global industry classification standard (GICS),” 2019. [Online]. Available: https://www.msci.com/gics
2019
Later among the works it cites.
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen, “Nemo: a toolkit for building ai applications using neural modules,” 2019
2019
Later among the works it cites.
O. Hrinchuk, M. Popova, and B. Ginsburg, “Correction of automatic speech recognition with transformer sequence-to-sequence model,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 05 2020, pp. 7074–7078
2020
Later among the works it cites.
lowerquality, Gentle Aligner , 2020 (accessed May 4th, 2020). [Online]. Available: https://lowerquality.com/gentle/
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
E. Pusateri, B. R. Ambati, E. Brooks, O. Plátek, D. McAllaster, and V. Nagesha, “A mostly data-driven approach to inverse text normalization,” in INTERSPEECH , 2017
2017
Cited alongside, same era.
M. Honnibal and I. Montani, “spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing,” 2017, to appear
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017. [Online]. Available: https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf
2017
Cited alongside, same era.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Estève, “TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation,” 2018
2018
Cited alongside, same era.
A. Akbik, D. Blythe, and R. Vollgraf, “Contextual string embeddings for sequence labeling,” in COLING 2018, 27th International Conference on Computational Linguistics , 2018, pp. 1638–1649
2018
Cited alongside, same era.
B. Nguyen, V. B. H. Nguyen, H. Nguyen, P. N. Phuong, T. Nguyen, Q. T. Do, and L. C. Mai, “Fast and accurate capitalization and punctuation for automatic speech recognition using transformer and chunk merging,” in 22nd Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques, O-COCOSDA 2019, Cebu, Philippines, October 25-27, 2019 . IEEE, 2019, pp. 1–5. [Online]. Available: https://doi.org/10.1109/O-COCOSDA46868.2019.9041202
2019
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” 2019
2019
Cited alongside, same era.
J. Wiseman, pyWebRTC , 2016 (accessed May 4th, 2020). [Online]. Available: https://github.com/wiseman/py-webrtcvad
2020
Later among the works it cites.
S. Global, “S&p capital iq,” www.capitaliq.com , 2012, accessed June 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
M. Sunkara, C. Shivade, S. Bodapati, and K. Kirchhoff, “Neural inverse text normalization,” 2021
2021
Closest in time.
J. Huang, O. Kuchaiev, P. O’Neill, V. Lavrukhin, J. Li, A. Flores, G. Kucsko, and B. Ginsburg, “Cross-language transfer learning, continuous learning, and domain adaptation for end-to-end automatic speech recognition,” in International Conference on Multimedia and Expo , 2021, to appear
2021
Closest in time.
G. Chen, S. Chai, G. Wang, J. Du, C. W. Wei-Qiang Zhang, D. Su, D. Povey, J. Trmal, J. Zhang, M. Ji, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Z. You, and Z. Yan, “GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,” submitted to INTERSPEECH , 2021
2021
Closest in time.