Fetching the paper…
Reading the bibliography…
As more and more information-rich data like video become available, utilizing multi-modal auxiliary information to enhance audio tasks has sparked widespread research interest.
W. H. Sumby and I. Pollack, “Visual contribution to speech intelligibility in noise,” Proc. JASA , 1954
1954
Earlier work this paper cites.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Proc. Nature , 1976
1976
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Kim et al. , “Hybrid ctc/attention architecture for end-to-end speech recognition,” Proc.JSTSP , 2017
2017
Earlier work this paper cites.
T. Afouras, J. S. Chung, and A. Zisserman, “LRS3-TED: a large-scale dataset for visual speech recognition,” arXiv preprint , 2018
2018
Earlier work this paper cites.
J. S. Chung, A. Nagrani, and A. Zisserman, “Voxceleb2: Deep speaker recognition,” arXiv preprint , 2018
2018
Earlier work this paper cites.
A. Ephrat, I. Mosseri, O. Lang et al. , “Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,” arXiv preprint , 2018
2018
Earlier work this paper cites.
S. Kim and F. Metze, “Dialog-context aware end-to-end speech recognition,” in Proc. SLT , 2018
2018
Earlier work this paper cites.
T. Makino, H. Liao, Y. Assael, B. Shillingford, B. Garcia, O. Braga, and O. Siohan, “Recurrent neural network transducer for audio-visual speech recognition,” in Proc. ASRU , 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
B. Xu, C. Lu, Y. Guo, and J. Wang, “Discriminative multi-modality speech recognition,” in Proc. CVPR , 2020
2020
Earlier work this paper cites.
J. Yu, S.-X. Zhang, J. Wu et al. , “Audio-visual recognition of overlapped speech for the LRS2 dataset,” in Proc. ICASSP , 2020
2020
Earlier work this paper cites.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in Proc. ICASSP , 2020
2020
Earlier work this paper cites.
T. Hori, N. Moritz, C. Hori, and J. Le Roux, “Transformer-based long-context end-to-end speech recognition,” in Proc. Interspeech , 2020
2020
Earlier work this paper cites.
P. Ma, S. Petridis, and M. Pantic, “End-to-end audio-visual speech recognition with conformers,” in Proc. ICASSP , 2021
2021
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” arXiv preprint , 2021
2021
Earlier work this paper cites.
G. Chen, S. Chai, G. Wang et al. , “Gigaspeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,” in Proc. Interspeech , 2021
2021
Cited alongside, same era.
C. Wang, M. Riviere, A. Lee et al. , “Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Proc. ACL , 2021
2021
Cited alongside, same era.
F.-J. Chang, J. Liu, M. Radfar, A. Mouchtaris, M. Omologo, A. Rastrow, and S. Kunzmann, “Context-aware transformer transducer for speech recognition,” in Proc. ASRU , 2021
2021
Cited alongside, same era.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” arXiv preprint , 2022
2022
Cited alongside, same era.
A. Haliassos, P. Ma, R. Mira, S. Petridis, and M. Pantic, “Jointly learning visual and auditory speech representations from raw data,” arXiv preprint , 2022
E. Lakomkin, C. Wu, Y. Fathullah, O. Kalinli, M. L. Seltzer, and C. Fuegen, “End-to-end speech recognition contextualization with large language models,” arXiv preprint , 2023
2023
Later among the works it cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux et al. , “LLaMA: Open and efficient foundation language models,” arXiv preprint , 2023
2023
Later among the works it cites.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu et al. , “Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,” https://vicuna.lmsys.org , 2023
2023
Later among the works it cites.
M. Cui, J. Kang, J. Deng et al. , “Towards effective and compact contextual representation for conformer transducer speech recognition systems,” arXiv preprint , 2023
2023
Later among the works it cites.
F. Yu, H. Wang, Z. Ma, and S. Zhang, “Hourglass-AVSR: Down-up sampling-based computational efficiency model for audio-visual speech recognition,” in Proc. ICASSP , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
G. Sun, C. Zhang, and P. C. Woodland, “Tree-constrained pointer generator with graph neural network encodings for contextual speech recognition,” in Proc. Interspeech , 2022
2022
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” in Proc. JSTSP , 2022
2022
Cited alongside, same era.
J. Li et al. , “Recent advances in end-to-end automatic speech recognition,” Proc. APSIPA , 2022
2022
Cited alongside, same era.
J. Wang, Z. Du, Q. Chen et al. , “LauraGPT: Listen, attend, understand, and regenerate audio with GPT,” in arXiv preprint , 2023
2023
Cited alongside, same era.
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu et al. , “On decoder-only architecture for speech-to-text and large language model integration,” in Proc. ASRU , 2023
2023
Cited alongside, same era.
S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua, “Next-gpt: Any-to-any multimodal llm,” in arXiv preprint , 2023
2023
Cited alongside, same era.
Y. Chu, J. Xu, X. Zhou et al. , “Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models,” arXiv preprint , 2023
2023
Cited alongside, same era.
2024
Closest in time.
H. Wang, P. Guo, P. Zhou, and L. Xie, “MLCA-AVSR: Multi-layer cross attention fusion based audio-visual speech recognition,” arXiv preprint , 2024
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin et al. , “Prompting large language models with speech recognition abilities,” in Proc. ICASSP , 2024
2024
Closest in time.
W. Yu, C. Tang, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “Connecting speech encoder and large language model for ASR,” in Proc. ICASSP , 2024
2024
Closest in time.
Z. Ma, G. Yang, Y. Yang, Z. Gao, J. Wang, Z. Du, F. Yu, Q. Chen, S. Zheng, S. Zhang et al. , “An embarrassingly simple approach for llm with strong asr capacity,” arXiv preprint , 2024
2024
Closest in time.
Z. Chen, H. Huang, Andrusenko et al. , “Salm: Speech-augmented language model with in-context learning for speech recognition and translation,” in Proc. ICASSP , 2024
2024
Closest in time.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” in Proc. ICLR , 2024
2024
Closest in time.
H. Wang, F. Yu, X. Shi, Y. Wang, and S. Zhang, “SlideSpeech: A large-scale slide-enriched audio-visual corpus,” in Proc. ICASPP , 2024
2024
Closest in time.
H. Wang, S. Kurita, S. Shimizu, and D. Kawahara, “SlideAVSR: A dataset of paper explanation videos for audio-visual speech recognition,” arXiv preprint , 2024
2024
Closest in time.
F. Yu, H. Wang, X. Shi, and S. Zhang, “LCB-NET: Long-context biasing for audio-visual speech recognition,” in Proc. ICASPP , 2024
2024
Closest in time.
X. Gong, Y. Wu, J. Li, S. Liu, R. Zhao, X. Chen, and Y. Qian, “Advanced long-content speech recognition with factorized neural transducer,” Proc. TASLP , 2024
2024
Closest in time.