Fetching the paper…
Reading the bibliography…
Text language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov et al., · 2015
Earlier work this paper cites.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin et al., · 2019
Earlier work this paper cites.
“CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al., · 2020
Earlier work this paper cites.
“Autoprompt: Eliciting knowledge from language models with automatically generated prompts,”
Taylor Shin et al., · 2020
Earlier work this paper cites.
“The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,”
Tu Anh Nguyen et al., · 2020
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
Earlier work this paper cites.
“Prefix-tuning: Optimizing continuous prompts for generation,”
Xiang Lisa Li and Percy Liang, · 2021
Earlier work this paper cites.
“The power of scale for parameter-efficient prompt tuning,”
Brian Lester, Rami Al-Rfou, and Noah Constant, · 2021
Earlier work this paper cites.
“On the effectiveness of adapter-based tuning for pretrained language model adaptation,”
Ruidan He et al., · 2021
Earlier work this paper cites.
“Finetuned language models are zero-shot learners,”
Jason Wei et al., · 2021
Earlier work this paper cites.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu wen Yang et al., · 2021
Earlier work this paper cites.
“ LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech,”
Solène Evain et al., · 2021
Earlier work this paper cites.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia et al., · 2021
Cited alongside, same era.
“Self-supervised speech representation learning: A review,”
Abdelrahman Mohamed et al., · 2022
Cited alongside, same era.
“An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks,”
Kai-Wei Chang et al., · 2022
Cited alongside, same era.
“P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks,”
Xiao Liu et al., · 2022
Cited alongside, same era.
“Adapterbias: Parameter-efficient token-dependent representation shift for adapters in nlp tasks,”
Chin-Lun Fu et al., · 2022
Cited alongside, same era.
“Slue: New benchmark tasks for spoken language understanding evaluation on natural speech,”
Suwon Shon et al., · 2022
“Audiogpt: Understanding and generating speech, music, sound, and talking head,”
Rongjie Huang et al., · 2023
Closest in time.
“Fleurs: Few-shot learning evaluation of universal representations of speech,”
Alexis Conneau et al., · 2023
Closest in time.
“Ml-superb: Multilingual speech universal performance benchmark,”
Jiatong Shi et al., · 2023
Closest in time.
“On the utility of self-supervised models for prosody-related tasks,”
Guan-Ting Lin et al., · 2023
Closest in time.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford et al., · 2023
Closest in time.
“Imagebind: One embedding space to bind them all,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Esb: A benchmark for multi-domain end-to-end speech recognition,”
Sanchit Gandhi, Patrick Von Platen, and Alexander M Rush, · 2022
Cited alongside, same era.
“Speechprompt v2: Prompt tuning for speech classification tasks,”
Kai-Wei Chang et al., · 2023
Cited alongside, same era.
“Speechgen: Unlocking the generative power of speech language models with prompts,”
Haibin Wu et al., · 2023
Cited alongside, same era.
“Gpt understands, too,”
Xiao Liu et al., · 2023
Cited alongside, same era.
“Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization,”
Puyuan Peng et al., · 2023
Cited alongside, same era.
“Exploring efficient-tuning methods in self-supervised speech models,”
Zih-Ching Chen et al., · 2023
Cited alongside, same era.
Rohit Girdhar et al., · 2023
Closest in time.
“Llama: Open and efficient foundation language models,”
Hugo Touvron et al., · 2023
Closest in time.
“Llama 2: Open foundation and fine-tuned chat models,”
Hugo Touvron et al., · 2023
Closest in time.
“Chatgpt (august 3 version),” Large language model, 2023
OpenAI, · 2023
Closest in time.
“Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,”
Yuan Gong et al., · 2023
Closest in time.
“Imagebind-llm: Multi-modality instruction tuning,”
Jiaming Han et al., · 2023
Closest in time.
“Llama-adapter v2: Parameter-efficient visual instruction model,”
Peng Gao et al., · 2023
Closest in time.
“Llama-adapter: Efficient fine-tuning of language models with zero-init attention,”
Renrui Zhang et al., · 2023
Closest in time.
“Can ChatGPT Detect Intent? Evaluating Large Language Models for Spoken Language Understanding,”
Mutian He and Philip N. Garner, · 2023
Closest in time.