Fetching the paper…
Reading the bibliography…
This paper explores the integration of Large Language Models (LLMs) into Automatic Speech Recognition (ASR) systems to improve transcription accuracy.
Graves, A., Fernández, S., Gomez, F., Schmidhuber, J.: Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In: Proceedings of the 23rd international conference on Machine learning. pp. 369–376 (2006)
2006
Earlier work this paper cites.
Graves, A., Mohamed, A.r., Hinton, G.: Speech recognition with deep recurrent neural networks. In: 2013 IEEE international conference on acoustics, speech and signal processing. pp. 6645–6649. Ieee (2013)
2013
Earlier work this paper cites.
Graves, A., Jaitly, N.: Towards end-to-end speech recognition with recurrent neural networks. In: International conference on machine learning. pp. 1764–1772. PMLR (2014)
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Chorowski, J.K., Bahdanau, D., Serdyuk, D., Cho, K., Bengio, Y.: Attention-based models for speech recognition. Advances in neural information processing systems 28
2015
Earlier work this paper cites.
Panayotov, V., Chen, G., Povey, D., Khudanpur, S.: Librispeech: an asr corpus based on public domain audio books. In: 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). pp. 5206–5210. IEEE (2015)
2015
Earlier work this paper cites.
Chan, W., Jaitly, N., Le, Q., Vinyals, O.: Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In: 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP). pp. 4960–4964. IEEE (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Bu, H., Du, J., Na, X., Wu, B., Zheng, H.: Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline. In: 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA). pp. 1–5. IEEE (2017)
2017
Earlier work this paper cites.
Kim, S., Hori, T., Watanabe, S.: Joint ctc-attention based end-to-end speech recognition using multi-task learning. In: 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). pp. 4835–4839. IEEE (2017)
2017
Earlier work this paper cites.
Tjandra, A., Sakti, S., Nakamura, S.: Listening while speaking: Speech chain by deep learning. In: 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). pp. 301–308. IEEE (2017)
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
Dong, L., Xu, S., Xu, B.: Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition. In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 5884–5888. IEEE (2018)
2018
Cited alongside, same era.
Kannan, A., Wu, Y., Nguyen, P., Sainath, T.N., Chen, Z., Prabhavalkar, R.: An analysis of incorporating an external language model into a sequence-to-sequence model. In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5828. IEEE (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Kubo, Y., Karita, S., Bacchiani, M.: Knowledge transfer from large-scale pretrained language models to end-to-end speech recognizers. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 8512–8516. IEEE (2022)
2022
Later among the works it cites.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35
2022
Later among the works it cites.
Peng, Y., Dalmia, S., Lane, I., Watanabe, S.: Branchformer: Parallel mlp-attention architectures to capture local and global context for speech recognition and understanding. In: International Conference on Machine Learning. pp. 17627–17643. PMLR (2022)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Shin, J., Lee, Y., Jung, K.: Effective sentence scoring method using bert for speech recognition. In: Asian Conference on Machine Learning. pp. 1081–1093. PMLR (2019)
2019
Cited alongside, same era.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Advances in neural information processing systems 33
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Chiu, S.H., Chen, B.: Innovative bert-based reranking language models for speech recognition. In: 2021 IEEE Spoken Language Technology Workshop (SLT). pp. 266–271. IEEE (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
Weiran, W., Chen, T., Sainath, T., Variani, E., Prabhavalkar, R., Huang, W.R., Ramabhadran, B., Gaur, N., Mavandadi, S., Peyser, C., Strohman, T., He, Y., Rybach, D.: Improving Rare Word Recognition with LM-aware MWER Training. In: Proc. Interspeech 2022. pp. 1031–1035 (2022). https://doi.org/10.21437/Interspeech.2022-10660
2022
Later among the works it cites.
Xu, L., Gu, Y., Kolehmainen, J., Khan, H., Gandhe, A., Rastrow, A., Stolcke, A., Bulyko, I.: Rescorebert: Discriminative speech recognition rescoring with bert. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 6117–6121. IEEE (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
OpenAI: Gpt-4 technical report (2023)
2023
Closest in time.