Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) have demonstrated commendable performance across a myriad of domains and tasks, existing LLMs still exhibit a palpable deficit in handling multimodal functionalities, especially for the Spoken Question Answering (SQA) task which necessitates precise alignment and deep interaction between speech and text features.
Brown T, Mann B, Ryder N, et al. Language models are few-shot learners. Advances in neural information processing systems, 2020, 33: 1877-1901
1901
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, et al. BLEU: a method for automatic evaluation of machine translation. Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 2002: 311-318
2002
Earlier work this paper cites.
Lin C Y. Rouge: A package for automatic evaluation of summaries. Text summarization branches out. 2004: 74-81
2004
Earlier work this paper cites.
Graves A, Fernández S, Gomez F, et al. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd international conference on Machine learning. 2006: 369-376
2006
Earlier work this paper cites.
Tur G, De Mori R. Spoken language understanding: Systems for extracting semantic information from speech. John Wiley & Sons, 2011
2011
Earlier work this paper cites.
Kearns J. Librivox: Free public domain audiobooks. Reference Reviews, 2014, 28(1): 7-8
2014
Earlier work this paper cites.
Panayotov V, Chen G, Povey D, et al. LibriSpeech: an asr corpus based on public domain audio books. IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015: 5206-5210
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Tapaswi M, Zhu Y, Stiefelhagen R, et al. MovieQA: Understanding stories in movies through question-answering. Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 4631-4640
2016
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in neural information processing systems, 2017, 30
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Serdyuk D, Wang Y, Fuegen C, et al. Towards end-to-end spoken language understanding. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018: 5754-5758
2018
Earlier work this paper cites.
Wang Y, Gales M J F, Knill K M, et al. Towards automatic assessment of spontaneous spoken English. Speech Communication, 2018, 104: 47-56
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Lee C H, Wang S M, Chang H C, et al. ODSQA: Open-domain spoken question answering dataset. IEEE Spoken Language Technology Workshop (SLT). IEEE, 2018: 949-956
2018
Earlier work this paper cites.
Malinin A. Uncertainty estimation in deep learning with application to spoken language assessment. University of Cambridge, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Zhang B, Sennrich R. Root mean square layer normalization. Advances in Neural Information Processing Systems, 2019, 32
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Zhao Z, Wang Y, Wang Y. Knowledge-Aware Bayesian Co-Attention for Multimodal Emotion Recognition. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Baevski A, Zhou Y, Mohamed A, et al. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems, 2020, 33: 12449-12460
2020
Cited alongside, same era.
Shazeer N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Hsu W N, Bolte B, Tsai Y H H, et al. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3451-3460
2021
Cited alongside, same era.
Wang C, Wu A, Gu J, et al. CoVoST 2 and massively multilingual speech translation. INTERSPEECH. 2021: 2247-2251
2021
Cited alongside, same era.
2022
Cited alongside, same era.
Agrawal B, Müller M, Choudhary S, et al. Tie your embeddings down: Cross-modal latent spaces for end-to-end spoken language understanding. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022: 7157-7161
2022
Cited alongside, same era.
Sun D, He Y, Han J. Using Auxiliary Tasks In Multimodal Fusion of wav2vec 2.0 And BERT for Multimodal Emotion Recognition. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5
2023
Closest in time.
Li F, Luo J, Wang L, et al. GCF2-Net: global-aware cross-modal feature fusion network for speech emotion recognition. Frontiers in Neuroscience, 2023, 17: 1183132
2023
Closest in time.
Laperrière G, Pelloin V, Rouvier M, et al. On the use of semantically-aligned speech representations for spoken language understanding. IEEE Spoken Language Technology Workshop (SLT). IEEE, 2023: 361-368
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision. International Conference on Machine Learning. PMLR, 2023: 28492-28518
2023
Closest in time.
2023
Closest in time.
Wang M, Han W, Shafran I, et al. SLM: Bridge the thin gap between speech and text foundation models. Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2023: 1-8
2023
Closest in time.
Deshmukh S, Elizalde B, Singh R, et al. Pengi: An audio language model for audio tasks. Advances in Neural Information Processing Systems, 2023, 36: 18090-18108
2023
Closest in time.
2023
Closest in time.
Taori R, Gulrajani I, Zhang T, et al. Stanford Alpaca: An instruction-following LLaMA model. 2023
2023
Closest in time.