Fetching the paper…
Reading the bibliography…
Recent studies have found that model performance has a smooth power-law relationship, or scaling laws, with training data and model size, for a wide range of problems.
T. Hori, C. Hori, S. Watanabe, and J. Hershey, “Minimum word error training of long short-term memory recurrent neural network language models for speech recognition,” in Proc. IEEE ICASSP , 2016, pp. 5990–5994
2016
Earlier work this paper cites.
Y. Xia, F. Tian, L. Wu, J. Lin, T. Qin, N. Yu, and T.-Y. Liu, “Deliberation networks: Sequence generation beyond one-pass decoding,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C.-C. Chiu, and A. Kannan, “Minimum word error rate training for attention-based sequence-to-sequence models,” in Proc. IEEE ICASSP , 2018, pp. 4839–4843
2018
Earlier work this paper cites.
T. N. Sainath, R. Pang, D. Rybach, Y. He, R. Prabhavalkar, W. Li, M. Visontai, Q. Liang, T. Strohman, Y. Wu, I. McGraw, and C.-C. Chiu, “Two-pass end-to-end speech recognition,” in Proc. Interspeech , 2019, pp. 2773–2777
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics , Minneapolis, Minnesota, Jun. 2019, pp. 4171–4186
2019
Earlier work this paper cites.
J. Shin, Y. Lee, and K. Jung, “Effective sentence scoring method using BERT for speech recognition,” in Asian Conference on Machine Learning . PMLR, 2019, pp. 1081–1093
2019
Earlier work this paper cites.
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang et al. , “Streaming end-to-end speech recognition for mobile devices,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6381–6385
2019
Earlier work this paper cites.
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, “Deliberation model based two-pass end-to-end speech recognition,” in Proc. IEEE ICASSP , 2020, pp. 7799–7803
2020
Earlier work this paper cites.
A. Gandhe and A. Rastrow, “Audio-attention discriminative language model for ASR rescoring,” in Proc. IEEE ICASSP , 2020, pp. 7944–7948
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
E. Variani, T. Chen, J. Apfel, B. Ramabhadran, S. Lee, and P. Moreno, “Neural oracle search on n-best hypotheses,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7824–7828
2020
Cited alongside, same era.
D. Fohr and I. Illina, “BERT-based semantic model for rescoring n-best speech recognition list,” in INTERSPEECH 2021 , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 12 104–12 113
2022
Later among the works it cites.
L. Xu, Y. Gu, J. Kolehmainen, H. Khan, A. Gandhe, A. Rastrow, A. Stolcke, and I. Bulyko, “RescoreBERT: Discriminative speech recognition rescoring with BERT,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6117–6121
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff, “Masked language model scoring,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, Jul. 2020, pp. 2699–2712
2020
Cited alongside, same era.
K. Hu, R. Pang, T. N. Sainath, and T. Strohman, “Transformer based deliberation for two-pass speech recognition,” in Proc. IEEE Spoken Language Technology Workshop , 2021, pp. 68–74
2021
Cited alongside, same era.
2021
Cited alongside, same era.
M. A. Gordon, K. Duh, and J. Kaplan, “Data and parameter scaling laws for neural machine translation,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 5915–5922
2021
Cited alongside, same era.
J. Droppo and O. Elibol, “Scaling laws for acoustic models,” arXiv preprint arXiv:2106.09488 , 2021
2021
Cited alongside, same era.
H. Futami, H. Inaguma, M. Mimura, S. Sakai, and T. Kawahara, “ASR rescoring and confidence estimation with ELECTRA,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2021, pp. 380–387
2021
Cited alongside, same era.
T.-W. Wu, I.-F. Chen, and A. Gandhe, “Learning to rank with BERT-based confidence models in ASR rescoring,” in Proc. Interspeech 2022 , 2022, pp. 1651–1655
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Hu, B. Li, and T. N. Sainath, “Scaling up deliberation for multilingual ASR,” in 2022 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2023, pp. 771–776
2023
Closest in time.