Fetching the paper…
Reading the bibliography…
Speech understanding is essential for interpreting the diverse forms of information embedded in spoken language, including linguistic, paralinguistic, and non-linguistic cues that are vital for effective human-computer interaction.
1904
Earlier work this paper cites.
R. Gray, “Vector quantization,” IEEE Assp Magazine , vol. 1, no. 2, pp. 4–29, 1984
1984
Earlier work this paper cites.
C. T. Hemphill, J. J. Godfrey, and G. R. Doddington, “The atis spoken language systems pilot corpus,” in Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990 , 1990
1990
Earlier work this paper cites.
D. A. Dahl, M. Bates, M. K. Brown, W. M. Fisher, K. Hunicke-Smith, D. S. Pallett, C. Pao, A. Rudnicky, and E. Shriberg, “Expanding the scope of the atis task: The atis-3 corpus,” in Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8-11, 1994 , 1994
1994
Earlier work this paper cites.
Q. Yang, J. Xu, W. Liu, Y. Chu, Z. Jiang, X. Zhou, Y. Leng, Y. Lv, Z. Zhao, C. Zhou, and J. Zhou, “AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , L.-W. Ku, A. Martins, and V. Srikumar, Eds., 2024, pp. 1979–1998
1998
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , P. Isabelle, E. Charniak, and D. Lin, Eds. Philadelphia, Pennsylvania, USA: Association for Computational Linguistics, Jul. 2002, pp. 311–318. [Online]. Available: https://aclanthology.org/P02-1040/
2002
Earlier work this paper cites.
P. Lamere et al. , “The cmu sphinx-4 speech recognition system,” Carnegie Mellon University, Pittsburgh, PA, USA, Tech. Rep , vol. 2003, 2003
2003
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text Summarization Branches Out: Proceedings of the ACL-04 Workshop , 2004, pp. 74–81
2004
Earlier work this paper cites.
W. Kraaij, T. Hain, M. Lincoln, and W. Post, “The ami meeting corpus,” in Proc. International Conference on Methods and Techniques in Behavioral Research , 2005, pp. 1–4
2005
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al. , “The ami meeting corpus: A pre-announcement,” in International workshop on machine learning for multimodal interaction . Springer, 2005, pp. 28–39
2005
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd International Conference on Machine Learning (ICML) , 2006, pp. 369–376
2006
Earlier work this paper cites.
J. S. B. Evans, “Dual-processing accounts of reasoning, judgment, and social cognition,” Annu. Rev. Psychol. , vol. 59, no. 1, pp. 255–278, 2008
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, pp. 335–359, 2008
2008
Earlier work this paper cites.
G. Tur and R. De Mori, Spoken language understanding: Systems for extracting semantic information from speech . John Wiley & Sons, 2011
2011
Earlier work this paper cites.
G. Tur and R. D. Mori, “Spoken language understanding: Systems for extracting semantic information from speech,” Wiley , 2011
2011
Earlier work this paper cites.
D. Kahneman, Thinking, fast and slow . macmillan, 2011
2011
Earlier work this paper cites.
D. Povey et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding . IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, B. Kingsbury et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” in IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 1. IEEE, 2012, pp. 30–42
2012
Earlier work this paper cites.
G. Tur, D. Hakkani-Tur, and S. Parthasarathy, “Asr error mitigation for improving spoken language understanding,” Computer Speech & Language , vol. 27, no. 3, pp. 664–684, 2013
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in International Conference on Machine Learning (ICML) , 2014, pp. 1764–1772
2014
Earlier work this paper cites.
G. Mesnil, Y. Dauphin, K. Yao, Y. Bengio, L. Deng, D. Hakkani-Tur, X. He, and L. Heck, “Using recurrent neural networks for slot filling in spoken language understanding,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 3, 2014, pp. 530–539
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
X. Yang and J. Liu, “Using word confusion networks for slot filling in spoken language understanding.” in Interspeech , 2015, pp. 1353–1357
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
R. Sarikaya, P. A. Crook, A. Marin, M. Jeong, J.-P. Robichaud, A. Celikyilmaz, Y.-B. Kim, A. Rochette, O. Z. Khan, X. Liu et al. , “An overview of end-to-end language understanding and dialog management for personal digital assistants,” in 2016 ieee spoken language technology workshop (slt) . IEEE, 2016, pp. 391–397
2016
Earlier work this paper cites.
F. Ladhak, M. Collins, and L. Zettlemoyer, “Lattice rescoring for speech-to-sql translation,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , 2016, pp. 2220–2225
2016
Earlier work this paper cites.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International conference on machine learning . PMLR, 2016, pp. 173–182
2016
Earlier work this paper cites.
D. Hakkani-Tür, G. Tur, A. Celikyilmaz, Y.-N. Chen, J. Gao, L. Deng, and Y.-Y. Wang, “Multi-domain joint semantic frame parsing using bi-directional rnn-lstm,” in INTERSPEECH , 2016, pp. 715–719
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to-end speech recognition using multi-task learning,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 4835–4839
2017
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,” in 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA) . IEEE, 2017, pp. 1–5
2017
Earlier work this paper cites.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “Must-c: a multilingual speech translation corpus,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics, 2019, pp. 2012–2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Haghani, A. Narayanan, M. Bacchiani, G. Chuang, N. Gaur, P. Moreno, R. Prabhavalkar, Z. Qu, and A. Waters, “From audio to semantics: Approaches to end-to-end spoken language understanding,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 720–726
2018
Earlier work this paper cites.
D. Serdyuk, Y. Wang, C. Fuegen, A. Kumar, B. Liu, and Y. Bengio, “Towards end-to-end spoken language understanding,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5754–5758
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Speech and Computer: 20th International Conference, SPECOM 2018, Leipzig, Germany, September 18–22, 2018, Proceedings 20 . Springer, 2018, pp. 198–208
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , vol. 13, no. 5, p. e0196391, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C.-H. Lee, S.-M. Wang, H.-C. Chang, and H.-Y. Lee, “Odsqa: Open-domain spoken question answering dataset,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 949–956
2018
Earlier work this paper cites.
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, “Speech model pretraining for end-to-end spoken language understanding,” in Proc. Interspeech , 2019, pp. 814–818
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Hou, Y. Shi, M. Ostendorf, M.-Y. Hwang, and L. Xie, “Region proposal network based small-footprint keyword spotting,” IEEE Signal Processing Letters , vol. 26, no. 10, pp. 1471–1475, 2019
2019
Earlier work this paper cites.
A. Coucke, M. Chlieh, T. Gisselbrecht, D. Leroy, M. Poumeyrol, and T. Lavril, “Efficient keyword spotting using dilated convolutions and gating,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6351–6355
2019
Earlier work this paper cites.
S. Zhu, Z. Zhao, T. Zhao, C. Zong, and K. Yu, “Catslu: The 1st chinese audio-textual spoken language understanding challenge,” in 2019 International Conference on Multimodal Interaction , 2019, pp. 521–525
2019
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
Q. Liu, Z. Chen, H. Li, M. Huang, Y. Lu, and K. Yu, “Modular end-to-end automatic speech recognition framework for acoustic-to-word model,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2174–2183, 2020
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in neural information processing systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Fan, J. Kang, L. Li, K. Li, H. Chen, S. Cheng, P. Zhang, Z. Zhou, Y. Cai, and D. Wang, “Cn-celeb: a challenging chinese speaker recognition dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7604–7608
2020
Earlier work this paper cites.
L. Martinez-Lucas, M. Abdelwahab, and C. Busso, “The msp-conversation corpus,” Interspeech 2020 , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” in IEEE Transactions on Audio, Speech, and Language Processing , 2021
2021
Cited alongside, same era.
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli et al. , “Prompting large language models with speech recognition abilities,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 351–13 355
2024
Closest in time.
2024
Closest in time.
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S.-w. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. Chang, S.-H. Lin, J. Liu, Y.-A. Chung, B.-H. Tseng, Y.-H. Fu, K. Lakhotia et al. , “Superb: Speech processing universal performance benchmark,” in Interspeech , 2021, pp. 1194–1198
2021
Cited alongside, same era.
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu, “W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2021, pp. 244–250
2021
Cited alongside, same era.
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, “Unsupervised speech recognition,” Advances in Neural Information Processing Systems , vol. 34, pp. 27 826–27 839, 2021
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
C. Wang, A. Wu, J. Gu, and J. Pino, “Covost 2 and massively multilingual speech translation.” in Interspeech , 2021, pp. 2247–2251
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Xue, Y. Liang, B. Mu, S. Zhang, M. Chen, Q. Chen, and L. Xie, “E-chat: Emotion-sensitive spoken dialogue system with large language models,” in 2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2024, pp. 586–590
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Yu, C. Tang, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “Connecting speech encoder and large language model for asr,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 637–12 641
2024
Closest in time.
Z. Chen, H. Huang, A. Andrusenko, O. Hrinchuk, K. C. Puvvada, J. Li, S. Ghosh, J. Balam, and B. Ginsburg, “Salm: Speech-augmented language model with in-context learning for speech recognition and translation,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 521–13 525
2024
Closest in time.
E. Lakomkin, C. Wu, Y. Fathullah, O. Kalinli, M. L. Seltzer, and C. Fuegen, “End-to-end speech recognition contextualization with large language models,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 406–12 410
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
N. B. Shankar, A. Johnson, C. Chance, H. Veeramani, and A. Alwan, “Coraal qa: A dataset and framework for open domain spontaneous speech question answering from long audio files,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 371–13 375
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Sclar, Y. Choi, Y. Tsvetkov, and A. Suhr, “Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=RIu5lyNXjT
2024
Closest in time.
2024
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Kumar, I. Thorbecke, S. Burdisso, E. Villatoro-Tello, M. K E, K. Hacioğlu, P. Rangappa, P. Motlicek, A. Ganapathiraju, and A. Stolcke, “Performance evaluation of slam-asr: The good, the bad, the ugly, and the way forward,” in 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW) , 2025, pp. 1–5
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
L.-C.-T. Xiaomi, “Mimo-audio: Audio language models are few-shot learners,” 2025. [Online]. Available: https://github.com/XiaomiMiMo/MiMo-Audio
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
K. Qian, Z. Chen, Y. Wu, S. Chen, J. Wu, M. Zeng, and X. Huang, “Speechformer: Redesigning the transformer architecture for end-to-end spoken language understanding,” in Proc. Interspeech , 2021, pp. 2097–2101
2097
Closest in time.