Fetching the paper…
Reading the bibliography…
Wav2Prompt is proposed which allows straightforward integration between spoken input and a text-based large language model (LLM).
Tools for the analysis of benchmark speech recognition tests, in: Proc. ICASSP
Pallet, D., Fisher, W., Fiscus, J., 1990 · 1990
Earlier work this paper cites.
A new algorithm for data compression
Gage, P., 1994 · 1994
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation, in: Proc. ACL
Papineni, K., Roukos, S., Ward, T., Zhu, W.J., 2002 · 2002
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks, in: Proc. ICML
Graves, A., Fernández, S., Gomez, F., Schmidhuber, J., 2006 · 2006
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books, in: Proc. ICASSP
Panayotov, V., Chen, G., Povey, D., Khudanpur, S., 2015 · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering, in: Proc. EMNLP
Yang, Y., Yih, W.t., Meek, C., 2015 · 2015
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit, in: Proc. Interspeech
Watanabe, S., Hori, T., Karita, S., Hayashi, T., Nishitoba, J., Unno, Y., Enrique Yalta Soplin, N., Heymann, J., Wiesner, M., Chen, N., Renduchintala, A., Ochiai, T., 2018 · 2018
Earlier work this paper cites.
Speech Model Pre-Training for End-to-End Spoken Language Understanding, in: Proc. Interspeech, pp. 814–818
Lugosch, L., Ravanelli, M., Ignoto, P., Tomar, V.S., Bengio, Y., 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners, in: Proc. NeurIPS
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 2020
Earlier work this paper cites.
CIF: Continuous integrate-and-fire for end-to-end speech recognition, in: Proc. ICASSP
Dong, L., Xu, B., 2020 · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented Transformer for speech recognition, in: Proc. Interspeech
Gulati, A., Qin, J., Chiu, C.C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., Pang, R., 2020 · 2020
Earlier work this paper cites.
Europarl-ST: A multilingual corpus for speech translation of parliamentary debates, in: Proc. ICASSP
Iranzo-Sánchez, J., Silvestre-Cerdà, J.A., Jorge, J., Roselló, N., Giménez, A., Sanchis, A., Civera, J., Juan, A., 2020 · 2020
Earlier work this paper cites.
Exploiting cloze-questions for few-shot text classification and natural language inference, in: Proc. EACL
Schick, T., Schütze, H., 2020 · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting knowledge from language models with automatically generated prompts, in: Proc. EMNLP
Shin, T., Razeghi, Y., IV, R.L.L., Wallace, E., Singh, S., 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing, in: Proc. EMNLP (Demos), pp. 38–45
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al., 2020 · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners, in: Zong, C., Xia, F., Li, W., Navigli, R. (Eds.), Proc. ACL/IJCNLP (1)
Gao, T., Fisch, A., Chen, D., 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning, in: Proc. EMNLP
Lester, B., Al-Rfou, R., Constant, N., 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, in: Proc. ICML, pp. 8748–8763
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al., 2021 · 2021
Cited alongside, same era.
Efficiently fusing pretrained acoustic and linguistic encoders for low-resource speech recognition
Yi, C., Zhou, S., Xu, B., 2021 · 2021
Cited alongside, same era.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks, in: Proc. ACL (2)
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., Tang, J., 2022 · 2022
Cited alongside, same era.
GPT understands, too
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., Tang, J., 2023 · 2023
Later among the works it cites.
Video-ChatGPT: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., Khan, F.S., 2023 · 2023
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Muennighoff, N., Wang, T., Sutawika, L., Roberts, A., Biderman, S., Scao, T.L., Bari, M.S., Shen, S., Yong, Z.X., Schoelkopf, H., Tang, X., Radev, D., Aji, A.F., Almubarak, K., Albanie, S., Alyafeai, Z., Webson, A., Raff, E., Raffel, C., 2023 · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision, in: Proc. ICML
Radford, A., Kim, J.W., Xu, T., Brockman, G., Mcleavey, C., Sutskever, I., 2023 · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al., 2022 · 2022
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization, in: Proc. ICLR
Victor, S., Albert, W., Colin, R., Stephen, B., Lintang, S., Zaid, A., Antoine, C., Arnaud, S., Arun, R., Manan, D., et al., 2022 · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E.H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., Fedus, W., 2022 · 2022
Cited alongside, same era.
AudioLM: A language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., et al., 2023 · 2023
Cited alongside, same era.
Chen, F., Han, M., Zhao, H., Zhang, Q., Shi, J., Xu, S., Xu, B., 2023 · 2023
Cited alongside, same era.
PaLM: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H.W., Sutton, C., Gehrmann, S., et al., 2023 · 2023
Cited alongside, same era.
Label-synchronous neural transducer for adaptable online E2E speech recognition
Deng, K., Woodland, P.C., 2023 · 2023
Cited alongside, same era.
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., Wei, F., 2023 · 2023
Later among the works it cites.
On decoder-only architecture for speech-to-text and large language model integration, in: Proc. ASRU
Wu, J., Gaur, Y., Chen, Z., Zhou, L., Zhu, Y., Wang, T., Li, J., Liu, S., Ren, B., Liu, L., Wu, Y., 2023 · 2023
Later among the works it cites.
SGP-TOD: Building task bots effortlessly via schema-guided LLM prompting, in: Proc. EMNLP (Findings)
Zhang, X., Peng, B., Li, K., Zhou, J., Meng, H., 2023 · 2023
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al., 2024 · 2024
Closest in time.
Prompting large language models with speech recognition abilities, in: Proc. ICASSP
Fathullah, Y., Wu, C., Lakomkin, E., Jia, J., Shangguan, Y., Li, K., Guo, J., Xiong, W., Mahadeokar, J., Kalinli, O., Fuegen, C., Seltzer, M., 2024a · 2024
Closest in time.
Listen, think, and understand, in: Proc. ICLR
Gong, Y., Luo, H., Liu, A.H., Karlinsky, L., Glass, J.R., 2024 · 2024
Closest in time.
Dynamic-Superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech, in: Proc. ICASSP
Huang, C.y., Lu, K.H., Wang, S.H., Hsiao, C.Y., Kuan, C.Y., Wu, H., Arora, S., Chang, K.W., Shi, J., Peng, Y., et al., 2024 · 2024
Closest in time.
SALMONN: Towards generic hearing abilities for large language models, in: Proc. ICLR
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., MA, Z., Zhang, C., 2024 · 2024
Closest in time.
A paradigm shift in machine translation: Boosting translation performance of large language models, in: Proc. ICLR
Xu, H., Kim, Y.J., Sharaf, A., Awadalla, H.H., 2024 · 2024
Closest in time.
Connecting speech encoder and large language model for ASR, in: Proc. ICASSP
Yu, W., Tang, C., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., Zhang, C., 2024 · 2024
Closest in time.