Fetching the paper…
Reading the bibliography…
In the realm of spoken language understanding (SLU), numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with transcribed speech instead of conventional written text.
“The ATIS spoken language systems pilot corpus,”
Charles T. Hemphill, John J. Godfrey, and George R. Doddington, · 1990
Earlier work this paper cites.
“The kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, et al., · 2011
Earlier work this paper cites.
“Minimum bayes risk decoding and system combination based on a recursion for edit distance,”
Haihua Xu, Daniel Povey, Lidia Mangu, and Jie Zhu, · 2011
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Lattice rnn: Recurrent neural networks over lattices,”
Faisal Ladhak, Ankur Gandhe, Markus Dreyer, Lambert Mathias, Ariya Rastrow, and Björn Hoffmeister, · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text,” 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang, · 2016
Earlier work this paper cites.
“Deep reinforcement learning from human preferences,”
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei, · 2017
Earlier work this paper cites.
“Neural confnet classification: Fully neural network based spoken utterance classification using word confusion networks,”
Ryo Masumura et al., · 2018
Earlier work this paper cites.
“Spoken squad: A study of mitigating the impact of speech recognition errors on listening comprehension,”
Chia-Hsuan Lee, Szu-Lin Wu, Chi-Liang Liu, and Hung-yi Lee, · 2018
Earlier work this paper cites.
“Adapting pretrained transformer to lattices for spoken language understanding,” 2020
Chao-Wei Huang and Yun-Nung Chen, · 2020
Earlier work this paper cites.
“Jointly encoding word confusion network and dialogue context with bert for spoken language understanding,” 2020
Chen Liu, Su Zhu, et al., · 2020
Cited alongside, same era.
“Slurp: A spoken language understanding resource package,” 2020
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser, · 2020
Cited alongside, same era.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Cited alongside, same era.
“N-best ASR transformer: Enhancing SLU performance using multiple ASR hypotheses,”
Karthik Ganesan, Pakhi Bamdev, Jaivarsan B, Amresh Venugopal, and Abhinav Tushar, · 2021
Cited alongside, same era.
“Datasets: A community library for natural language processing,”
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, et al., · 2021
Cited alongside, same era.
“Training language models to follow instructions with human feedback,”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, et al., · 2022
Later among the works it cites.
“DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering,”
Guan-Ting Lin, Yung-Sung Chuang, et al., · 2022
Later among the works it cites.
“Promptchainer: Chaining large language model prompts through visual programming,”
Tongshuang Wu, Ellen Jiang, Aaron Donsbach, Jeff Gray, Alejandra Molina, Michael Terry, and Carrie J Cai, · 2022
Later among the works it cites.
Kai-Wei Chang, Wei-Cheng Tseng, Shang-Wen Li, and Hung-yi Lee, · 2022
Later among the works it cites.
“Voicebox: Text-guided multilingual universal speech generation at scale,”
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al., · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Voice2series: Reprogramming acoustic models for time series classification,”
Chao-Han Huck Yang, Yun-Yun Tsai, and Pin-Yu Chen, · 2021
Cited alongside, same era.
“Crosslingual generalization through multitask finetuning,”
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, et al., · 2022
Cited alongside, same era.
“Bloom: A 176b-parameter open-access multilingual language model,”
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al., · 2022
Cited alongside, same era.
“Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks,”
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, et al., · 2022
Cited alongside, same era.
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Later among the works it cites.
“Can chatgpt detect intent? evaluating large language models for spoken language understanding,” 2023
Mutian He and Philip N. Garner, · 2023
Later among the works it cites.
“Speechprompt v2: Prompt tuning for speech classification tasks,”
Kai-Wei Chang, Yu-Kai Wang, Hua Shen, Iu-thing Kang, Wei-Cheng Tseng, Shang-Wen Li, and Hung-yi Lee, · 2023
Later among the works it cites.
“Speechgen: Unlocking the generative power of speech language models with prompts,”
Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu, and Hung-yi Lee, · 2023
Later among the works it cites.