Fetching the paper…
Reading the bibliography…
Audio-LLM introduces audio modality into a large language model (LLM) to enable a powerful LLM to recognize, understand, and generate audio.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in International Conference on Machine Learning, ICML 2006 , ser. ACM International Conference Proceeding Series, W. W. Cohen and A. W. Moore, Eds., vol. 148. ACM, 2006, pp. 369–376
2006
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Annual Conference on Neural Information Processing Systems, NeurIPS 2017 , I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, Eds., 2017, pp. 5998–6008
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline,” in Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems and Assessment, O-COCOSDA 2017 . IEEE, 2017, pp. 1–5
2017
Earlier work this paper cites.
A. Fan, M. Lewis, and Y. N. Dauphin, “Hierarchical neural story generation,” in Annual Meeting of the Association for Computational Linguistics, ACL 2018 , I. Gurevych and Y. Miyao, Eds. Association for Computational Linguistics, 2018, pp. 889–898
2018
Earlier work this paper cites.
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi, “The curious case of neural text degeneration,” in International Conference on Learning Representations, ICLR 2020 . OpenReview.net, 2020
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Annual Conference of the International Speech Communication Association, INTERSPEECH 2020 , H. Meng, B. Xu, and T. F. Zheng, Eds. ISCA, 2020, pp. 5036–5040
2020
Earlier work this paper cites.
B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, D. Wu, and Z. Peng, “WENETSPEECH: A 10000+ hours multi-domain mandarin corpus for speech recognition,” in International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022 . IEEE, 2022, pp. 6182–6186
2022
Earlier work this paper cites.
OpenAI, “GPT-4 technical report,” CoRR , vol. abs/2303.08774, 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif et al. , “Palm: Scaling language modeling with pathways,” J. Mach. Learn. Res. , vol. 24, pp. 240:1–240:113, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
C. H. Yang, Y. Gu, Y. Liu, S. Ghosh, I. Bulyko, and A. Stolcke, “Generative speech recognition error correction with large language models and task-activating prompting,” in Automatic Speech Recognition and Understanding Workshop, ASRU 2023 . IEEE, 2023, pp. 1–8
2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, “Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,” in Conference on Empirical Methods in Natural Language Processing, EMNLP 2023 , H. Bouamor, J. Pino, and K. Bali, Eds. Association for Computational Linguistics, 2023, pp. 15 757–15 773
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Wang, W. Han, I. Shafran, Z. Wu, C. Chiu, Y. Cao, N. Chen, Y. Zhang, H. Soltau, P. K. Rubenstein, L. Zilka, D. Yu, G. Pundak, N. Siddhartha, J. Schalkwyk, and Y. Wu, “SLM: bridge the thin gap between speech and text foundation models,” in Automatic Speech Recognition and Understanding Workshop, ASRU 2023 . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Later among the works it cites.
C. Chen, Y. Hu, C. H. Yang, S. M. Siniscalchi, P. Chen, and C. E. Siong, “Hyporadise: An open baseline for generative speech recognition with large language models,” in Annual Conference on Neural Information Processing Systems, NeurIPS 2023 , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning, ICML 2023 , ser. Proceedings of Machine Learning Research, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, Eds., vol. 202. PMLR, 2023, pp. 28 492–28 518
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Hu, C. Chen, C.-H. H. Yang, R. Li, C. Zhang, P.-Y. Chen, and E. S. Chng, “Large language models are efficient learners of noise-robust speech recognition,” in International Conference on Learning Representations, ICLR 2024 . OpenReview.net, 2024
2024
Closest in time.