Fetching the paper…
Reading the bibliography…
Integrating speech into LLM (speech-LLM) has gaining increased attention recently.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Machine Learning, Proceedings of the Twenty-Third International Conference (ICML 2006), Pittsburgh, Pennsylvania, USA, June 25-29, 2006 , ser. ACM International Conference Proceeding Series, vol. 148. ACM, 2006, pp. 369–376
2006
Earlier work this paper cites.
T. B. Brown, B. Mann et al. , “Language models are few-shot learners,” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , 2020
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in 21st Annual Conference of the International Speech Communication Association, Interspeech 2020, Virtual Event, Shanghai, China, October 25-29, 2020 . ISCA, 2020, pp. 5036–5040
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Earlier work this paper cites.
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised cross-lingual representation learning at scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , 2020, pp. 8440–8451
2020
Earlier work this paper cites.
R. Fan, W. Chu, P. Chang, and J. Xiao, “CASS-NAT: CTC alignment-based single step non-autoregressive transformer for speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2021, Toronto, ON, Canada, June 6-11, 2021 . IEEE, 2021, pp. 5889–5893
2021
Earlier work this paper cites.
M. Gaido, M. Cettolo, M. Negri, and M. Turchi, “Ctc-based compression for direct speech translation,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021 . Association for Computational Linguistics, 2021, pp. 690–696
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
J. Ao, R. Wang, L. Zhou, C. Wang et al. , “SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,” in Proc. ACL , 2022, pp. 5723–5738
2022
Earlier work this paper cites.
J. Zhao, H. Yang, G. Haffari, and E. Shareghi, “M-adapter: Modality adaptation for end-to-end speech-to-text translation,” in Proc. Interspeech , 2022, pp. 111–115
2022
Earlier work this paper cites.
Z. Chen, Y. Zhang, A. Rosenberg, B. Ramabhadran, P. J. Moreno, A. Bapna, and H. Zen, “MAESTRO: Matched speech text representations through modality matching,” in Proc. Interspeech , 2022, pp. 4093–4097
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 , 2022
2022
Earlier work this paper cites.
J. Achiam, S. Adler et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023
Earlier work this paper cites.
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi, “Self-instruct: Aligning language models with self-generated instructions,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 13 484–13 508
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu, “On decoder-only architecture for speech-to-text and large language model integration,” in IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023, Taipei, Taiwan, December 16-20, 2023 . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Trans. Mach. Learn. Res. , vol. 2023, 2023
2023
Cited alongside, same era.
Y. Fathullah, C. Wu, E. Lakomkin, K. Li, J. Jia, S. Yuan, J. Mahadeokar, O. Kalinli, C. Fuegen, and M. Seltzer, “Audiochatllama: Towards general-purpose speech abilities for llms,” in North American Chapter of the Association for Computational Linguistics , 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:265150137
2023
Cited alongside, same era.
J. Pan, J. Wu, Y. Gaur, S. Sivasankaran, Z. Chen, S. Liu, and J. Li, “COSMIC: data efficient instruction-tuning for speech in-context learning,” Interspeech , 2024
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli et al. , “Prompting large language models with speech recognition abilities,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 351–13 355
2024
Closest in time.
2024
Closest in time.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
R. Fan, W. Chu, P. Chang, and A. Alwan, “A CTC alignment-based non-autoregressive transformer for end-to-end automatic speech recognition,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 31, pp. 1436–1448, 2023
2023
Cited alongside, same era.
2024
Cited alongside, same era.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
S. Hu, L. Zhou et al. , “Wavllm: Towards robust and adaptive speech large language model,” in Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024 . Association for Computational Linguistics, 2024, pp. 4552–4572
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Fan, N. B. Shankar, and A. Alwan, “Unienc-cassnat: An encoder-only non-autoregressive asr for speech ssl models,” IEEE Signal Processing Letters , vol. 31, pp. 711–715, 2024
2024
Closest in time.
R. Zhao, J. Li, R. Fan, and M. Post, “CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation,” SLT , 2024
2024
Closest in time.
Z. Zhang, S. Chen, L. Zhou, Y. Wu, S. Ren, S. Liu, Z. Yao, X. Gong, L. Dai, J. Li et al. , “SpeechLM: Enhanced speech pre-training with unpaired textual data,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.
2024
Closest in time.
K. Deng, G. Sun, and P. Woodland, “Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , Apr. 2025, pp. 6940–6956
2025
Closest in time.