Fetching the paper…
Reading the bibliography…
Incorporating speech understanding capabilities into pretrained large-language models has become a vital research direction (SpeechLLM).
2012
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
W. Chan et al. , “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP . IEEE, 2016, pp. 4960–4964
2016
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Brown et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Interspeech , 2020
2020
Earlier work this paper cites.
C. Raffel et al. , “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, no. 140, pp. 1–67, 2020. [Online]. Available: http://jmlr.org/papers/v21/20-074.html
2020
Earlier work this paper cites.
Q. Zhang et al. , “Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,” in ICASSP . IEEE, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
J. Kahn et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP , 2020, pp. 7669–7673, https://github.com/facebookresearch/libri-light
2020
Earlier work this paper cites.
R. Ardila et al. , “Common voice: A massively-multilingual speech corpus,” in LREC , 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
J.-B. Alayrac et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” 2022
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass, “Joint audio and speech understanding,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Wu et al. , “On decoder-only architecture for speech-to-text and large language model integration,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
2023
Cited alongside, same era.
M. Wang et al. , “Slm: Bridge the thin gap between speech and text foundation models,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
J. Bai et al. , “Qwen technical report,” arXiv preprint arXiv:2309.16609 , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
V. Srivastav et al. , “Open automatic speech recognition leaderboard,” https://huggingface.co/spaces/huggingface.co/spaces/open-asr-leaderboard/leaderboard , 2023
2023
Later among the works it cites.
A. Conneau et al. , “Fleurs: Few-shot learning evaluation of universal representations of speech,” in SLT . IEEE, 2023, pp. 798–805
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
Z. Chen et al. , “Salm: Speech-augmented language model with in-context learning for speech recognition and translation,” in ICASSP . IEEE, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
E. Tsunoo, H. Futami, Y. Kashiwagi, S. Arora, and S. Watanabe, “Decoder-only architecture for streaming end-to-end speech recognition,” 2024
2024
Closest in time.
2024
Closest in time.
V. Noroozi, Z. Chen et al. , “Instruction data generation and unsupervised adaptation for speech language models,” in Interspeech , 2024
2024
Closest in time.
“Canary-1b model,” https://huggingface.co/nvidia/canary-1b , accessed: 2024-03-11
2024
Closest in time.
P. Zhang, G. Zeng, T. Wang, and W. Lu, “Tinyllama: An open-source small language model,” 2024
2024
Closest in time.