Fetching the paper…
Reading the bibliography…
Recent studies leverage large language models with multi-tasking capabilities, using natural language prompts to guide the model's behavior and surpassing performance of task-specific models.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Multitask learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
IEMOCAP: interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008 · 2008
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
Tie your embeddings down: Cross-modal latent spaces for end-to-end spoken language understanding
Bhuvan Agrawal, Markus Müller, Martin Radfar, Samridhi Choudhary, Athanasios Mouchtaris, and Siegfried Kunzmann. 2020 · 2011
Earlier work this paper cites.
Acquisition of ordinal words using weakly supervised NMF
Vincent Renkens, Steven Janssens, Bart Ons, Jort F. Gemmeke, and Hugo Van hamme. 2014 · 2014
Earlier work this paper cites.
Online multitask learning for machine translation quality estimation
José Guilherme Camargo de Souza, Matteo Negri, Elisa Ricci, and Marco Turchi. 2015 · 2015
Earlier work this paper cites.
A unified perspective on multi-domain and multi-task learning
Yongxin Yang and Timothy M. Hospedales. 2015 · 2015
Earlier work this paper cites.
Multi-task learning for speech recognition: an overview
Gueorgui Pironkov, Stéphane Dupont, and Thierry Dutoit. 2016 · 2016
Earlier work this paper cites.
Freesound datasets: A platform for the creation of open audio datasets
Eduardo Fonseca, Jordi Pons, Xavier Favory, Frederic Font, Dmitry Bogdanov, Andres Ferraro, Sergio Oramas, Alastair Porter, and Xavier Serra. 2017 · 2017
Earlier work this paper cites.
Joint CTC-attention based end-to-end speech recognition using multi-task learning
Suyoun Kim, Takaaki Hori, and Shinji Watanabe. 2017 · 2017
Earlier work this paper cites.
Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition
Shubham Toshniwal, Hao Tang, Liang Lu, and Karen Livescu. 2017 · 2017
Earlier work this paper cites.
Tied multitask learning for neural speech translation
Antonios Anastasopoulos and David Chiang. 2018 · 2018
Earlier work this paper cites.
Voxforge
Ken MacLean. 2018 · 2018
Earlier work this paper cites.
Spoken language understanding on the edge
Alaa Saade, Alice Coucke, Alexandre Caulier, Joseph Dureau, Adrien Ball, Théodore Bluche, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, and Maël Primet. 2018 · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition
Pete Warden. 2018 · 2018
Earlier work this paper cites.
Towards multimodal sarcasm detection (an _obviously_ perfect paper)
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. 2019 · 2019
Earlier work this paper cites.
Speech model pre-training for end-to-end SLU
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio. 2019 · 2019
Earlier work this paper cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton. 2019 · 2019
Earlier work this paper cites.
Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech
Andreas Nautsch, Xin Wang, Nicholas W. D. Evans, Tomi H. Kinnunen, Ville Vestman, Massimiliano Todisco, Héctor Delgado, Md. Sahidullah, Junichi Yamagishi, and Kong Aik Lee. 2021 · 2019
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019 · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Different absorption from the same sharing: Sifted multi-task learning for fake news detection
Lianwei Wu, Yuan Rao, Haolin Jin, Ambreen Nazir, and Ling Sun. 2019 · 2019
Earlier work this paper cites.
A neural multi-task learning framework to jointly model medical named entity recognition and normalization
Sendong Zhao, Ting Liu, Sicheng Zhao, and Fei Wang. 2019 · 2019
Earlier work this paper cites.
Pintext: A multitask text embedding system in pinterest
Jinfeng Zhuang and Yu Liu. 2019 · 2019
Cited alongside, same era.
Accentdb: A database of non-native english accents to assist neural speech recognition
Afroz Ahamad, Ankit Anand, and Pranesh Bhargava. 2020 · 2020
Cited alongside, same era.
SLURP: A spoken language understanding resource package
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser. 2020 · 2020
Cited alongside, same era.
Unsupervised pre-training for voice activation
Aliaksei Kolesau and Dmitrij Šešok. 2020 · 2020
Cited alongside, same era.
MAD-X: an adapter-based framework for multi-task cross-lingual transfer
Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Cited alongside, same era.
Multi-task learning for speaker verification and voice trigger detection
Siddharth Sigtia, Erik Marchi, Sachin Kajarekar, Devang Naik, and John Bridle. 2020 · 2020
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022 · 2022
Later among the works it cites.
A multimodal corpus for emotion recognition in sarcasm
Anupama Ray, Shubham Mishra, Apoorva Nunna, and Pushpak Bhattacharyya. 2022 · 2022
Later among the works it cites.
Universal paralinguistic speech representations using self-supervised conformers
Joel Shor, Aren Jansen, Wei Han, Daniel S. Park, and Yu Zhang. 2022 · 2022
Later among the works it cites.
STOP: A dataset for Spoken Task Oriented Semantic Parsing
Paden Tomasello, Akshat Shrivastava, Daniel Lazar, Po-Chun Hsu, Duc Le, Adithya Sagar, Ali Elkahky, Jade Copet, Wei-Ning Hsu, Yossef Mordechay, Robin Algayres, Tu Anh Nguyen, Emmanuel Dupoux, Luke Zettlemoyer, and Abdelrahman Mohamed. 2022 · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improving end-to-end speech-to-intent classification with reptile
Yusheng Tian and Philip John Gorinski. 2020 · 2020
Cited alongside, same era.
Speech emotion recognition with multi-task learning
Xingyu Cai, Jiahong Yuan, Renjie Zheng, Liang Huang, and Kenneth Church. 2021 · 2021
Cited alongside, same era.
Speechstew: Simply mix all available speech recognition data to train one large neural network
William Chan, Daniel S. Park, Chris A. Lee, Yu Zhang, Quoc V. Le, and Mohammad Norouzi. 2021 · 2021
Cited alongside, same era.
Multi-task learning in natural language processing: An overview
Shijie Chen, Yu Zhang, and Qiang Yang. 2021 · 2021
Cited alongside, same era.
Building and benchmarking an arabic speech commands dataset for small-footprint keyword spotting
Abdulkader Ghandoura, Farouk Hjabo, and Oumayma Al Dakkak. 2021 · 2021
Cited alongside, same era.
Voxceleb enrichment for age and gender recognition
Khaled Hechmi, Trung Ngo Trong, Ville Hautamäki, and Tomi Kinnunen. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Tandem multitask training of speaker diarisation and speech recognition for meeting transcription
Xianrui Zheng, Chao Zhang, and Philip C. Woodland. 2022 · 2022
Later among the works it cites.
Siddhant Arora, Hayato Futami, Shih-Lun Wu, Jessica Huynh, Yifan Peng, Yosuke Kashiwagi, Emiru Tsunoo, Brian Yan, and Shinji Watanabe. 2023 · 2023
Closest in time.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023 · 2023
Closest in time.
SpeechPrompt v2: Prompt tuning for speech classification tasks
Kai-Wei Chang, Yu-Kai Wang, Hua Shen, Iu-thing Kang, Wei-Cheng Tseng, Shang-Wen Li, and Hung-yi Lee. 2023 · 2023
Closest in time.
Zhehuai Chen, He Huang, Andrei Andrusenko, Oleksii Hrinchuk, Krishna C Puvvada, Jason Li, Subhankar Ghosh, Jagadeesh Balam, and Boris Ginsburg. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023 · 2023
Closest in time.
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023 · 2023
Closest in time.
Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, Roshan S. Sharma, Shinji Watanabe, Bhiksha Ramakrishnan, Shady Shehata, and Hung-yi Lee. 2023 · 2023
Closest in time.
Instruction-following speech recognition
Cheng-I Jeff Lai, Zhiyun Lu, Liangliang Cao, and Ruoming Pang. 2023 · 2023
Closest in time.
End-to-end speech recognition contextualization with large language models
Egor Lakomkin, Chunyang Wu, Yassir Fathullah, Ozlem Kalinli, Michael L Seltzer, and Christian Fuegen. 2023 · 2023
Closest in time.
Dailytalk: Spoken dialogue dataset for conversational text-to-speech
Keon Lee, Kyumin Park, and Daeyoung Kim. 2023 · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Closest in time.
Soumi Maiti, Yifan Peng, Shukjae Choi, Jee-weon Jung, Xuankai Chang, and Shinji Watanabe. 2023 · 2023
Closest in time.
Whisper-slu: Extending a pretrained speech-to-text transformer for low resource spoken language understanding
Quentin Meeus, Marie-Francine Moens, and Hugo Van Hamme. 2023 · 2023
Closest in time.
End-to-end speech recognition: A survey
Rohit Prabhavalkar, Takaaki Hori, Tara N. Sainath, Ralf Schlüter, and Shinji Watanabe. 2023 · 2023
Closest in time.
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023 · 2023
Closest in time.
SALMONN: towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023 · 2023
Closest in time.
Lauragpt: Listen, attend, understand, and regenerate audio with GPT
Jiaming Wang, Zhihao Du, Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang, Siqi Zheng, Chang Zhou, Zhijie Yan, and Shiliang Zhang. 2023 · 2023
Closest in time.
Connecting speech encoder and large language model for asr
Wenyi Yu, Changli Tang, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023 · 2023
Closest in time.
Google USM: scaling automatic speech recognition beyond 100 languages
Yu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, Zhong Meng, Ke Hu, Andrew Rosenberg, Rohit Prabhavalkar, Daniel S. Park, Parisa Haghani, Jason Riesa, Ginger Perng, Hagen Soltau, Trevor Strohman, Bhuvana Ramabhadran, Tara N. Sainath, Pedro J. Moreno, Chung-Cheng Chiu, Johan Schalkwyk, Françoise Beaufays, and Yonghui Wu. 2023 · 2023
Closest in time.