Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have gained considerable attention for Artificial Intelligence Generated Content (AIGC), particularly with the emergence of ChatGPT.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al., 2020 · 1901
Earlier work this paper cites.
Covost 2 and massively multilingual speech-to-text translation
Wang, C., Wu, A., Pino, J., 2020 · 2007
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R.L., Wallace, E., Singh, S., 2020 · 2010
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books, in: 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE. pp. 5206–5210
Panayotov, V., Chen, G., Povey, D., Khudanpur, S., 2015 · 2015
Earlier work this paper cites.
The lj speech dataset
Ito, K., Johnson, L., 2017 · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A.v.d., Li, Y., Vinyals, O., 2018 · 2018
Earlier work this paper cites.
Adversarial reprogramming of neural networks, in: ICLR (Poster), OpenReview.net
Elsayed, G.F., Goodfellow, I.J., Sohl-Dickstein, J., 2019 · 2019
Earlier work this paper cites.
Language models as knowledge bases?, in: EMNLP/IJCNLP (1), Association for Computational Linguistics. pp. 2463–2473
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P.S.H., Bakhtin, A., Wu, Y., Miller, A.H., 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., Auli, M., 2020 · 2020
Earlier work this paper cites.
Gpt-3: Its nature, scope, limits, and consequences
Floridi, L., Chiriatti, M., 2020 · 2020
Earlier work this paper cites.
How can we know what language models know
Jiang, Z., Xu, F.F., Araki, J., Neubig, G., 2020 · 2020
Earlier work this paper cites.
Transfer learning without knowing: Reprogramming black-box machine learning models with scarce data and limited resources, in: ICML, PMLR. pp. 9614–9624
Tsai, Y., Chen, P., Ho, T., 2020 · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D.A., Adeli, E., Altman, R.B., Arora, S., von Arx, S., Bernstein, M.S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N.S., Chen, A.S., Creel, K., Davis, J.Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N.D., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D.E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P.W., Krass, M.S., Krishna, R., Kuditipudi, R., et al., 2021 · 2021
Earlier work this paper cites.
Speechnet: A universal modularized model for speech processing tasks
Chen, Y., Chi, P., Yang, S., Chang, K., Lin, J., Huang, S., Liu, D., Liu, C., Lee, C., Lee, H., 2021 · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W., Bolte, B., Tsai, Y.H., Lakhotia, K., Salakhutdinov, R., Mohamed, A., 2021 · 2021
Earlier work this paper cites.
On generative spoken language modeling from raw audio
Lakhotia, K., Kharitonov, E., Hsu, W.N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.A., Copet, J., Baevski, A., Mohamed, A., et al., 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning, in: EMNLP (1), Association for Computational Linguistics. pp. 3045–3059
Lester, B., Al-Rfou, R., Constant, N., 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation, in: ACL/IJCNLP (1), Association for Computational Linguistics. pp. 4582–4597
Li, X.L., Liang, P., 2021 · 2021
Cited alongside, same era.
Prompt programming for large language models: Beyond the few-shot paradigm, in: CHI Extended Abstracts, ACM. pp. 314:1–314:7
Reynolds, L., McDonell, K., 2021 · 2021
Cited alongside, same era.
Yen, H., Ku, P.J., Yang, C.H.H., Hu, H., Siniscalchi, S.M., Chen, P.Y., Tsao, Y., 2021 · 2021
Cited alongside, same era.
Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing, in: ACL (1), Association for Computational Linguistics. pp. 5723–5738
Black-box tuning for language-model-as-a-service, in: ICML, PMLR. pp. 20841–20855
Sun, T., Shao, Y., Qian, H., Huang, X., Qiu, X., 2022 · 2022
Later among the works it cites.
SUPERB-SG: enhanced speech processing universal performance benchmark for semantic and generative capabilities, in: ACL (1), Association for Computational Linguistics. pp. 8479–8492
Tsai, H., Chang, H., Huang, W., Huang, Z., Lakhotia, K., Yang, S., Dong, S., Liu, A.T., Lai, C., Shi, J., Chang, X., Hall, P., Chen, H., Li, S., Watanabe, S., Mohamed, A., Lee, H., 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., Zhou, D., 2022 · 2022
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., Tagliasacchi, M., 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ao, J., Wang, R., Zhou, L., Wang, C., Ren, S., Wu, Y., Liu, S., Ko, T., Li, Q., Zhang, Y., Wei, Z., Qian, Y., Li, J., Wei, F., 2022 · 2022
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., Zeghidour, N., 2022 · 2022
Cited alongside, same era.
An exploration of prompt tuning on generative spoken language model for speech processing tasks, in: INTERSPEECH, ISCA. pp. 5005–5009
Chang, K., Tseng, W., Li, S., Lee, H., 2022 · 2022
Cited alongside, same era.
Model reprogramming: Resource-efficient cross-domain machine learning
Chen, P., 2022 · 2022
Cited alongside, same era.
Speech-to-speech translation for A real-world unwritten language
Chen, P., Tran, K., Yang, Y., Du, J., Kao, J., Chung, Y., Tomasello, P., Duquenne, P., Schwenk, H., Gong, H., Inaguma, H., Popuri, S., Wang, C., Pino, J.M., Hsu, W., Lee, A., 2022 · 2022
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., Adi, Y., 2022 · 2022
Cited alongside, same era.
Superb @ SLT 2022: Challenge on generalization and efficiency of self-supervised speech representation learning, in: SLT, IEEE. pp. 1096–1103
Feng, T., Dong, S.A., Yeh, C., Yang, S., Lin, T., Shi, J., Chang, K., Huang, Z., Wu, H., Chang, X., Watanabe, S., Mohamed, A., Li, S., Lee, H., 2022 · 2022
Cited alongside, same era.
Wavprompt: Towards few-shot spoken language understanding with frozen language models, in: INTERSPEECH, ISCA. pp. 2738–2742
Gao, H., Ni, J., Qian, K., Zhang, Y., Chang, S., Hasegawa-Johnson, M., 2022 · 2022
Cited alongside, same era.
Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H., Ba, J., 2022 · 2022
Later among the works it cites.
Speechprompt v2: Prompt tuning for speech classification tasks
Chang, K., Wang, Y., Shen, H., Kang, I., Tseng, W., Li, S., Lee, H., 2023 · 2023
Closest in time.
Textually pretrained speech language models
Hassid, M., Remez, T., Nguyen, T.A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Defossez, A., Synnaeve, G., Dupoux, E., et al., 2023 · 2023
Closest in time.
Low-resource music genre classification with cross-modal neural model reprogramming, in: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE. pp. 1–5
Hung, Y.N., Yang, C.H.H., Chen, P.Y., Lerch, A., 2023 · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., Zeghidour, N., 2023 · 2023
Closest in time.
Lms with a voice: Spoken language modeling beyond speech tokens
Nachmani, E., Levkovitch, A., Salazar, J., Asawaroengchai, C., Mariooryad, S., Skerry-Ryan, R., Ramanovich, M.T., 2023 · 2023
Closest in time.
Prompting the hidden talent of web-scale speech models for zero-shot task generalization
Peng, P., Yan, B., Watanabe, S., Harwath, D., 2023 · 2023
Closest in time.
Radhakrishnan, S., Yang, C.H.H., Khan, S.A., Kiani, N.A., Gomez-Cabrero, D., Tegner, J.N., 2023 · 2023
Closest in time.
A prompt pattern catalog to enhance prompt engineering with chatgpt
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., Schmidt, D.C., 2023 · 2023
Closest in time.
Yang, C.H.H., Li, B., Zhang, Y., Chen, N., Prabhavalkar, R., Sainath, T.N., Strohman, T., 2023 · 2023
Closest in time.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Zhang, Z., Zhou, L., Wang, C., Chen, S., Wu, Y., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., Wei, F., 2023 · 2023
Closest in time.