Fetching the paper…
Reading the bibliography…
We introduce Speech ReaLLM, a new ASR architecture that marries "decoder-only" ASR with the RNN-T to make multimodal LLM architectures capable of real-time streaming.
1906
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
2012
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proceedings of ICASSP , 2015
2015
Earlier work this paper cites.
R. Prabhavalkar, K. Rao, T. N. Sainath, and Others, “A comparison of sequence-to-sequence models for speech recognition,” in Proceedings of Interspeech , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C.-C. Chiu* and C. Raffel*, “Monotonic chunkwise attention,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=Hko85plCW
2018
Earlier work this paper cites.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Y. Shi, Y. Wang, C. Wu, and Others, “Emformer: Efficient memory transformer based acoustic model for low latency streaming speech recognition,” in Proceedings of ICASSP , 2021
2021
Cited alongside, same era.
J. Wu, Y. Gaur, Z. Chen, and Others, “On decoder-only architecture for speech-to-text and large language model integration,” in Proceedings of ASRU , 2023, pp. 1–8
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, and Others, “Robust speech recognition via Large-Scale weak supervision,” in Proceedings of ICML , 2023
2023
Later among the works it cites.
L. Barrault et al. , “Seamlessm4t: Massively multilingual & multimodal machine translation,” ArXiv , 2023
2023
Later among the works it cites.
X. Ma, A. Sun, S. Ouyang, H. Inaguma, and P. Tomasello, “Efficient monotonic multihead attention,” 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Shi, C. Wu, D. Wang, and Others, “Streaming transformer transducer based speech recognition using Non-Causal convolution,” in Proceedings of ICASSP , 2022
2022
Cited alongside, same era.
J. Bach, “Synthetic sentience: Can artificial intelligence become conscious?” in 37th Chaos Communication Congress (37C3): Unlocked , 2023. [Online]. Available: https://www.youtube.com/watch?v=Ms96Py8p8Jg
2023
Cited alongside, same era.
Cited in the paper.
2023
Later among the works it cites.