Fetching the paper…
Reading the bibliography…
CTC compressor can be an effective approach to integrate audio encoders to decoder-only models, which has gained growing interest for different speech applications.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” 2012, arXiv:1211.3711
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Estève, “Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks,” in
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in
2018
Earlier work this paper cites.
B. Zoph, C.-C. Chiu, D. S. Park, E. D. Cubuk, Q. V. Le, W. Chan, and Y. Zhang, “SpecAugment: A Simple Augmentation Method for Automatic Speech Recognition,” in
2019
Earlier work this paper cites.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in
2020
Cited alongside, same era.
M. Gaido, M. Cettolo, M. Negri, and M. Turchi, “CTC-based compression for direct speech translation,” in
2021
Cited alongside, same era.
Z. Meng, N. Kanda, Y. Gaur, S. Parthasarathy, E. Sun, L. Lu, X. Chen, J. Li, and Y. Gong, “Internal language model training for domain-adaptive end-to-end speech recognition,” in
2021
Cited alongside, same era.
T. N. Sainath, R. Prabhavalkar, A. Bapna, Y. Zhang, Z. Huo, Z. Chen, B. Li, W. Wang, and T. Strohman, “JOIST: A joint speech and text streaming model for ASR,” in
2022
Cited alongside, same era.
S. Thomas, B. Kingsbury, G. Saon, and H. J. Kuo, “Integrating text inputs for training and adapting RNN transducer ASR models,” in
Google, “Gemini: A family of highly capable multimodal models,” 2024, arXiv:2312.11805
2024
Closest in time.
Meta AI, “The Llama 3 herd of models,” 2024, arXiv:2407.21783
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli, C. Fuegen, and M. Seltzer, “Prompting large language models with speech recognition abilities,” in
2024
Closest in time.
E. Tsunoo, H. Futami, Y. Kashiwagi, S. Arora, and S. Watanabe, “Decoder-only architecture for streaming end-to-end speech recognition,” in
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
Z. Chen, Y. Zhang, A. Rosenberg, B. Ramabhadran, P. J. Moreno, A. Bapna, and H. Zen, “MAESTRO: matched speech text representations through modality matching,” in
2022
Cited alongside, same era.
OpenAI, “GPT-4 technical report,” 2023, arXiv:2303.08774
2023
Cited alongside, same era.
Meta AI, “Llama 2: Open foundation and fine-tuned chat models,” 2023, arXiv:2307.09288
2023
Cited alongside, same era.
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu, “On decoder-only architecture for speech-to-text and large language model integration,” in
2023
Cited alongside, same era.
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, “Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,” in
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
S. Ling, Y. Hu, S. Qian, G. Ye, Y. Qian, Y. Gong, E. Lin, and M. Zeng, “Adapting large language model with speech for fully formatted end-to-end speech recognition,” in
2024
Closest in time.
2024
Closest in time.
Z. Zhang, S. Chen, L. Zhou, Y. Wu, S. Ren, S. Liu, Z. Yao, X. Gong, L. Dai, J. Li, and F. Wei, “Speechlm: Enhanced speech pre-training with unpaired textual data,”
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, K. Li, J. Jia, Y. Shangguan, J. Mahadeokar, O. Kalinli, C. Fuegen, and M. Seltzer, “Audiochatllama: Towards general-purpose speech abilities for llms,” in
2024
Closest in time.
J. Su, M. H. M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,”
2024
Closest in time.