Fetching the paper…
Reading the bibliography…
We investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering.
D.-C. Lyu, T. P. Tan, C. E. Siong, and H. Li, “Seame: a mandarin-english code-switching speech corpus in south-east asia,” in Interspeech , 2010
2010
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in ACCV , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” 2017
2017
Earlier work this paper cites.
R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze, “How2: a large-scale dataset for multimodal language understanding,” 2018
2018
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “Espnet: End-to-end speech processing toolkit,” in Interspeech , 2018
2018
Earlier work this paper cites.
A. C. Kocabiyikoglu, L. Besacier, and O. Kraif, “Augmenting librispeech with French translations: A multimodal corpus for direct speech translation evaluation,” in LREC , 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Earlier work this paper cites.
A. Miech, D. Zhukov, J. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,” in ICCV , 2019
2019
Earlier work this paper cites.
B. Wu, W. Chen, Y. Fan, Y. Zhang, J. Hou, J. Liu, and T. Zhang, “Tencent ml-images: A large-scale multi-label image database for visual representation learning,” IEEE Access , 2019
2019
Earlier work this paper cites.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
Y. Chung, W. Weng, S. Tong, and J. R. Glass, “Towards unsupervised speech-to-text translation,” in ICASSP , 2019
2019
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in NeurIPS , 2020
2020
Earlier work this paper cites.
A. Baevski, H. Zhou, A. rahman Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , 2020
2020
Cited alongside, same era.
C. Wang, A. Wu, and J. Pino, “Covost 2: A massively multilingual speech-to-text translation corpus,” 2020
2020
Cited alongside, same era.
C. Escolano, M. R. Costa-jussà, J. A. R. Fonollosa, and C. Segura, “Enabling zero-shot multilingual spoken language translation with language-specific encoders and decoders,” ASRU , 2020
2020
Cited alongside, same era.
C. Wang, Y. Tang, X. Ma, A. Wu, D. Okhonko, and J. Pino, “Fairseq S2T: Fast speech-to-text modeling with fairseq,” in AACL: System Demonstrations , 2020
2020
Cited alongside, same era.
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill et al. , “On the opportunities and risks of foundation models,” ArXiv preprint , 2021
K.-W. Chang, W.-C. Tseng, S.-W. Li, and H.-y. Lee, “An exploration of prompt tuning on generative spoken language model for speech processing tasks,” in Interspeech , 2022
2022
Later among the works it cites.
J. ho Kim, J.-S. Heo, H. seo Shin, C. Lim, and H. jin Yu, “Integrated parameter-efficient tuning for general-purpose audio models,” ArXiv , 2022
2022
Later among the works it cites.
J. Xue, P. Wang, J. Li, and E. Sun, “A weakly-supervised streaming multilingual speech model with truly zero-shot capability,” ArXiv , 2022
2022
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” ArXiv preprint , 2022
2022
Later among the works it cites.
V. Gabeur, P. H. Seo, A. Nagrani, C. Sun, K. Alahari, and C. Schmid, “Avatar: Unconstrained audiovisual speech recognition,” in Interspeech , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Cited alongside, same era.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, and E. Dupoux, “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , 2021
2021
Cited alongside, same era.
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in ICLR , 2022
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou, “Chain of thought prompting elicits reasoning in large language models,” in NeurIPS , 2022
2022
Cited alongside, same era.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” IJCV , no. 9, 2022
2022
Cited alongside, same era.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. hsin Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus, “Emergent abilities of large language models,” TMLR , 2022
2022
Cited alongside, same era.
H. Gao, J. Ni, K. Qian, Y. Zhang, S. Chang, and M. Hasegawa-Johnson, “Wavprompt: Towards few-shot spoken language understanding with frozen language models,” in Interspeech , 2022
2022
Cited alongside, same era.
2022
Later among the works it cites.
G. I. Winata, A. F. Aji, Z.-X. Yong, and T. Solorio, “The decades progress on code-switching research in nlp: A systematic survey on trends and challenges,” ArXiv preprint , 2022
2022
Later among the works it cites.
H. Lovenia, S. Cahyawijaya, G. Winata, P. Xu, Y. Xu, Z. Liu, R. Frieske, T. Yu, W. Dai, E. J. Barezi, Q. Chen, X. Ma, B. Shi, and P. Fung, “ASCEND: A spontaneous Chinese-English dataset for code-switching in multi-turn conversation,” in LREC , 2022
2022
Later among the works it cites.
T. Nguyen, N. Tran, L. Deng, T. G. da Silva, M. Radzihovsky, R. Hsiao, H. Mason, S. Braun, E. McDermott, D. Can, P. Swietojanski, L. Verwimp, S. Oyman, T. Arvizo, H. Silovsky, A. Ghoshal, M. J. Martel, B. R. Ambati, and M. Ali, “Optimizing bilingual neural transducer with synthetic code-switching text generation,” ArXiv , 2022
2022
Later among the works it cites.
C. Wang, H. Inaguma, P.-J. Chen, I. Kulikov, Y. Tang, W.-N. Hsu, M. Auli, and J. M. Pino, “Simple and effective unsupervised speech translation,” ArXiv , 2022
2022
Later among the works it cites.
P.-A. Duquenne, H. Gong, B. Sagot, and H. Schwenk, “T-modules: Translation modules for zero-shot cross-modal machine translation,” in EMNLP , 2022
2022
Later among the works it cites.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , no. 9, 2023
2023
Closest in time.
A. Zeng, M. Attarian, brian ichter, K. M. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. S. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. Florence, “Socratic models: Composing zero-shot multimodal reasoning with language,” in ICLR , 2023
2023
Closest in time.