Fetching the paper…
Reading the bibliography…
We introduce MERaLiON-AudioLLM (Multimodal Empathetic Reasoning and Learning in One Network), the first speech-text model tailored for Singapore's multilingual and multicultural landscape.
BLEU: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
VoxCeleb: A large-scale speaker identification dataset
A. Nagrani, J. S. Chung, and A. Zisserman · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
S. Elfwing, E. Uchibe, and K. Doya · 2018
Earlier work this paper cites.
Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension
C.-H. Lee, S.-L. Wu, C.-L. Liu, and H.-y. Lee · 2018
Earlier work this paper cites.
Building the Singapore English National Speech Corpus
J. X. Koh, A. Mislan, K. Khoo, B. Ang, W. Ang, C. Ng, and Y.-Y. Tan · 2019
Earlier work this paper cites.
SpecAugment: A simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Earlier work this paper cites.
MELD: A multimodal multi-party dataset for emotion recognition in conversations
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea · 2019
Earlier work this paper cites.
Common Voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber · 2020
Earlier work this paper cites.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Earlier work this paper cites.
Earnings-21: A practical benchmark for ASR in the wild
M. Rio, N. Delworth, R. Westerman, M. Huang, N. Bhandari, J. Palakapilly, Q. McNamara, J. Dong, P. Żelasko, and M. Jette · 2021
Earlier work this paper cites.
CoVoST 2 and massively multilingual speech translation
C. Wang, A. Wu, J. Gu, and J. Pino · 2021
Earlier work this paper cites.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, and et al · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Earlier work this paper cites.
Earnings-22: A practical benchmark for accents in the wild
M. Rio, H. Peter, Q. McNamara, C. Miller, and S. Chandra · 2022
Earlier work this paper cites.
streaming
T. M. M. Team · 2022
Earlier work this paper cites.
Press release on Singapore’s National Multimodal Large Language Model Programme, 2023
A*STAR · 2023
Earlier work this paper cites.
Qwen technical report
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, and et al · 2023
Earlier work this paper cites.
BEATs: Audio pre-training with acoustic tokenizers
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei · 2023
Earlier work this paper cites.
Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models
Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Pengi: An audio language model for audio tasks
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang · 2023
Cited alongside, same era.
Joint audio and speech understanding
Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass · 2023
Cited alongside, same era.
Gaussian error linear units (GELUs)
D. Hendrycks and K. Gimpel · 2023
Cited alongside, same era.
ConvMLP: Hierarchical convolutional MLPs for vision
J. Li, A. Hassani, S. Walton, and H. Shi · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2023
Cited alongside, same era.
AudioPaLM: A large language model that can speak and listen
P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. de Chaumont Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov, H. Muckenhirn, D. Padfield, J. Qin, D. Rozenberg, T. Sainath, J. Schalkwyk, M. Sharifi, M. T. Ramanovich, M. Tagliasacchi, A. Tudor, M. Velimirović, D. Vincent, J. Yu, Y. Wang, V. Zayats, N. Zeghidour, Y. Zhang, Z. Zhang, L. Zilka, and C. Frank · 2023
GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities
S. Ghosh, S. Kumar, A. Seth, C. K. R. Evuru, U. Tyagi, S. Sakshi, O. Nieto, R. Duraiswami, and D. Manocha · 2024
Closest in time.
Listen, think, and understand
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. Glass · 2024
Closest in time.
Distilling an end-to-end voice assistant without instruction training data
W. Held, E. Li, M. Ryan, W. Shi, Y. Zhang, and D. Yang · 2024
Closest in time.
WavLLM: Towards robust and adaptive speech large language model
S. Hu, L. Zhou, S. Liu, S. Chen, L. Meng, H. Hao, J. Pan, X. Liu, J. Li, S. Sivasankaran, L. Liu, and F. Wei · 2024
Closest in time.
WavChat: A survey of spoken dialogue models
S. Ji, Y. Chen, M. Fang, J. Zuo, J. Lu, H. Wang, Z. Jiang, L. Zhou, S. Liu, X. Cheng, X. Yang, Z. Wang, Q. Yang, J. Li, Y. Jiang, J. He, Y. Chu, J. Xu, and Z. Zhao · 2024
Closest in time.
Frozen large language models can perceive paralinguistic aspects of speech
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
SLUE phase-2: A benchmark suite of diverse spoken language understanding tasks
S. Shon, S. Arora, C.-J. Lin, A. Pasad, F. Wu, R. S. Sharma, W.-L. Wu, H.-y. Lee, K. Livescu, and S. Watanabe · 2023
Cited alongside, same era.
Stanford Alpaca: An instruction-following LLaMA model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Cited alongside, same era.
OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants, 2023
Teknium · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. R. Stone, P. Albert, A. Almahairi, and et al · 2023
Cited alongside, same era.
VioLA: Unified codec language models for speech recognition, synthesis, and translation
T. Wang, L. Zhou, Z. Zhang, Y. Wu, S. Liu, Y. Gaur, Z. Chen, J. Li, and F. Wei · 2023
Cited alongside, same era.
On decoder-only architecture for speech-to-text and large language model integration
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu · 2023
Cited alongside, same era.
W. Kang, J. Jia, C. Wu, W. Zhou, E. Lakomkin, Y. Gaur, L. Sari, S. Kim, K. Li, J. Mahadeokar, and O. Kalinli · 2024
Closest in time.
Audio Flamingo: A novel audio language model with few-shot learning and dialogue abilities
Z. Kong, A. Goel, R. Badlani, W. Ping, R. Valle, and B. Catanzaro · 2024
Closest in time.
DeSTA: Enhancing speech language models through descriptive speech-text alignment
K.-H. Lu, Z. Chen, S.-W. Fu, H. Huang, B. Ginsburg, Y.-C. F. Wang, and H. yi Lee · 2024
Closest in time.
An embarrassingly simple approach for llm with strong asr capacity
Z. Ma, G. Yang, Y. Yang, Z. Gao, J. Wang, Z. Du, F. Yu, Q. Chen, S. Zheng, S. Zhang, and X. Chen · 2024
Closest in time.
Meralion-speechencoder: Towards a speech foundation model for singapore and beyond, 2024
MERaLiON Team · 2024
Closest in time.
Large language models: A survey
S. Minaee, T. Mikolov, N. Nikzad, M. A. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Closest in time.
Spirit LM: Interleaved spoken and written language model
T. A. Nguyen, B. Muller, B. Yu, M. R. Costa-jussa, M. Elbayad, S. Popuri, C. Ropers, P.-A. Duquenne, R. Algayres, R. Mavlyutov, I. Gat, M. Williamson, G. Synnaeve, J. Pino, B. Sagot, and E. Dupoux · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, J. Ferret, P. Liu, P. Tafti, A. Friesen, M. Casbon, S. Ramos, R. Kumar, C. L. Lan, S. Jerome, A. Tsitsulin, N. Vieillard, P. Stanczyk, S. Girgin, N. Momchev, M. Hoffman, S. Thakoor, J.-B. Grill, B. Neyshabur, O. Bachem, A. Walton, A. Severyn, A. Parrish, A. Ahmad, A. Hutchison, A. Abdagic, A. Carl, A. Shen, A. Brock, A. Coenen, A. Laforge, A. Paterson, B. Bastian, B. Piot, B. Wu, B. Royal, C. Chen, C. Kumar, C. Perry, C. Welty, C. A. Choquette-Choo, D. Sinopalnikov, D. Weinberger, D. Vijaykumar, D. Rogozińska, D. Herbison, E. Bandy, E. Wang, E. Noland, E. Moreira, E. Senter, E. Eltyshev, F. Visin, G. Rasskin, G. Wei, G. Cameron, G. Martins, H. Hashemi, H. Klimczak-Plucińska, H. Batra, H. Dhand, I. Nardini, J. Mein, J. Zhou, J. Svensson, J. Stanway, J. Chan, J. P. Zhou, J. Carrasqueira, J. Iljazi, J. Becker, J. Fernandez, J. van Amersfoort, J. Gordon, J. Lipschultz, J. Newlan, J. yeong Ji, K. Mohamed, K. Badola, K. Black, K. Millican, K. McDonell, K. Nguyen, K. Sodhia, K. Greene, L. L. Sjoesund, L. Usui, L. Sifre, L. Heuermann, L. Lago, L. McNealus, L. B. Soares, L. Kilpatrick, L. Dixon, L. Martins, M. Reid, M. Singh, M. Iverson, M. Görner, M. Velloso, M. Wirth, M. Davidow, M. Miller, M. Rahtz, M. Watson, M. Risdal, M. Kazemi, M. Moynihan, M. Zhang, M. Kahng, M. Park, M. Rahman, M. Khatwani, N. Dao, N. Bardoliwalla, N. Devanathan, N. Dumai, N. Chauhan, O. Wahltinez, P. Botarda, P. Barnes, P. Barham, P. Michel, P. Jin, P. Georgiev, P. Culliton, P. Kuppala, R. Comanescu, R. Merhej, R. Jana, R. A. Rokni, R. Agarwal, R. Mullins, S. Saadat, S. M. Carthy, S. Cogan, S. Perrin, S. M. R. Arnold, S. Krause, S. Dai, S. Garg, S. Sheth, S. Ronstrom, S. Chan, T. Jordan, T. Yu, T. Eccles, T. Hennigan, T. Kocisky, T. Doshi, V. Jain, V. Yadav, V. Meshram, V. Dharmadhikari, W. Barkley, W. Wei, W. Ye, W. Han, W. Kwon, X. Xu, Z. Shen, Z. Gong, Z. Wei, V. Cotruta, P. Kirk, A. Rao, M. Giang, L. Peran, T. Warkentin, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, D. Sculley, J. Banks, A. Dragan, S. Petrov, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, S. Borgeaud, N. Fiedel, A. Joulin, K. Kenealy, R. Dadashi, and A. Andreev · 2024
Closest in time.
Audiobench: A universal benchmark for audio large language models
B. Wang, X. Zou, G. Lin, S. Sun, Z. Liu, W. Zhang, Z. Liu, A. Aw, and N. F. Chen · 2024
Closest in time.
Mini-Omni: Language models can hear, talk while thinking in streaming
Z. Xie and C. Wu · 2024
Closest in time.
Comparing discrete and continuous space LLMs for speech recognition
Y. Xu, S.-X. Zhang, J. Yu, Z. Wu, and D. Yu · 2024
Closest in time.
Connecting speech encoder and large language model for ASR
W. Yu, C. Tang, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang · 2024
Closest in time.
MoWE-Audio: Multitask audiollms with mixture of weak encoders
W. Zhang, S. Sun, B. Wang, X. Zou, Z. Liu, Y. He, G. Lin, N. F. Chen, and A. T. Aw · 2024
Closest in time.
Advancing singlish understanding: Bridging the gap with datasets and multimodal models
B. Wang, X. Zou, S. Sun, W. Zhang, Y. He, Z. Liu, C. Wei, N. F. Chen, and A. Aw · 2025
Closest in time.