Fetching the paper…
Reading the bibliography…
Spoken language models (SLMs) that integrate speech with large language models (LMs) rely on modality adapters (MAs) to map the output of speech encoders to a representation that is understandable to the decoder LM.
F. Hill, R. Reichart, and A. Korhonen, “SimLex-999: Evaluating semantic models with (genuine) similarity estimation,” Computational Linguistics , vol. 41, 2015
2015
Earlier work this paper cites.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi,” in Interspeech 2017 , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. R. Mortensen, S. Dalmia, and P. Littell, “Epitran: Precision G2P for many languages,” in Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , 2018
2018
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of the Twelfth Language Resources and Evaluation Conference , 2020. [Online]. Available: https://aclanthology.org/2020.lrec-1.520/
2020
Earlier work this paper cites.
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2021
2021
Earlier work this paper cites.
Y.-A. Chung, Y. Belinkov, and J. Glass, “Similarity analysis of self-supervised speech representations,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Earlier work this paper cites.
Z. Fan, M. Li, S. Zhou, and B. Xu, “Exploring wav2vec 2.0 on speaker verification and language identification,” in Interspeech 2021 , 2021, pp. 1509–1513
2021
Earlier work this paper cites.
B. van Niekerk, L. Nortje, M. Baas, and H. Kamper, “Analyzing speaker information in self-supervised models to improve zero-resource speech processing,” in Interspeech 2021 , 2021, pp. 1554–1558
2021
Earlier work this paper cites.
Z.-Y. Dou and G. Neubig, “Word alignment by fine-tuning embeddings on parallel corpora,” in Conference of the European Chapter of the Association for Computational Linguistics (EACL) , 2021
2021
Earlier work this paper cites.
D. Merkx, S. Frank, and M. Ernestus, “SpokenSTS,” 2021. [Online]. Available: https://doi.org/10.17026/dans-z48-3ev6
2021
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
S. Chen, Y. Wu, C. Wang, S. Liu, Z. Chen, P. Wang, G. Liu, J. Li, J. Wu, X. Yu, and F. Wei, “Why does self-supervised learning for speech recognition benefit speaker recognition?” in Interspeech 2022 , 2022, pp. 3699–3703
2022
Cited alongside, same era.
M. Bartelds and M. Wieling, “Quantifying language variation acoustically with few resources,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2022, pp. 3735–3741
2022
Cited alongside, same era.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio pre-training with acoustic tokenizers,” in Proceedings of the 40th International Conference on Machine Learning , 2023
2023
Later among the works it cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” 2023
2023
Later among the works it cites.
2024
Later among the works it cites.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
E. Ahn and E. Chodroff, “VoxCommunis: A corpus for cross-linguistic phonetic analysis,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference , 2022
2022
Cited alongside, same era.
N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzmán, and A. Fan, “The FLORES-101 evaluation benchmark for low-resource and multilingual machine translation,” Transactions of the Association for Computational Linguistics , vol. 10, pp. 522–538, 2022
2022
Cited alongside, same era.
Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass, “Joint audio and speech understanding,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
N. Prakash and R. K.-W. Lee, “Layered bias: Interpreting bias in pretrained large language models,” in Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP . Association for Computational Linguistics, 2023, pp. 284–295
2023
Cited alongside, same era.
B. M. Abdullah, M. M. Shaik, B. Möbius, and D. Klakow, “An information-theoretic analysis of self-supervised discrete representations of speech,” in Interspeech 2023 , 2023, pp. 2883–2887
2023
Cited alongside, same era.
Nostalgebraist, “Interpreting GPT: The logit lens.” [Online]. Available: https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens
Cited in the paper.
A. Pasad, C.-M. Chien, S. Settle, and K. Livescu, “What do self-supervised speech models know about words?” Transactions of the Association for Computational Linguistics , vol. 12, pp. 372–391, 2024
2024
Later among the works it cites.
C. Wendler, V. Veselovsky, G. Monea, and R. West, “Do llamas work in English? on the latent language of multilingual transformers,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, 2024
2024
Later among the works it cites.
T. Tang, W. Luo, H. Huang, D. Zhang, X. Wang, X. Zhao, F. Wei, and J.-R. Wen, “Language-specific neurons: The key to multilingual capabilities in large language models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association for Computational Linguistics, 2024, pp. 5701–5715
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.