Fetching the paper…
Reading the bibliography…
In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction.
“Nemo: a toolkit for building ai applications using neural modules,”
Oleksii Kuchaiev et al., · 2019
Earlier work this paper cites.
“autochord: Automatic Chord Recognition Library and Chord Visualization App,”
Christopher John Bayron, · 2021
Earlier work this paper cites.
“The Zero Resource Speech Challenge 2021: Spoken Language Modelling,”
Ewan Dunbar, Mathieu Bernard, Nicolas Hamilakis, Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Eugene Kharitonov, and Emmanuel Dupoux, · 2021
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models,”
Jason Wei et al., · 2022
Earlier work this paper cites.
“Emergent abilities of large language models,”
Barret Zoph et al., · 2022
Earlier work this paper cites.
“Large language models are zero-shot reasoners,”
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa, · 2022
Earlier work this paper cites.
“Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,”
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch, · 2022
Earlier work this paper cites.
“Titanet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context,”
Nithin Rao Koluguri, Taejin Park, and Boris Ginsburg, · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback,”
Long Ouyang et al., · 2022
Earlier work this paper cites.
“On the planning abilities of large language models-a critical investigation,”
Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati, · 2023
Earlier work this paper cites.
“Large language models can self-improve,”
Jiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han, · 2023
Earlier work this paper cites.
“Vipergpt: Visual inference via python execution for reasoning,”
Dídac Surís, Sachit Menon, and Carl Vondrick, · 2023
Earlier work this paper cites.
“Creator: Tool creation for disentangling abstract and concrete reasoning of large language models,”
Cheng Qian, Chi Han, Yi Fung, Yujia Qin, Zhiyuan Liu, and Heng Ji, · 2023
Earlier work this paper cites.
“Api-bank: A comprehensive benchmark for tool-augmented llms,”
Minghao Li et al., · 2023
Earlier work this paper cites.
“Gorilla: Large language model connected with massive apis,”
Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E Gonzalez, · 2023
Earlier work this paper cites.
“Can large language models be an alternative to human evaluations?,”
Cheng-Han Chiang and Hung-Yi Lee, · 2023
Earlier work this paper cites.
“A closer look into using large language models for automatic evaluation,”
Cheng-Han Chiang and Hung-yi Lee, · 2023
Earlier work this paper cites.
“Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,”
Yunfei Chu et al., · 2023
Earlier work this paper cites.
“Joint audio and speech understanding,”
Yuan Gong, Alexander H Liu, Hongyin Luo, Leonid Karlinsky, and James Glass, · 2023
Earlier work this paper cites.
“Chatgpt: Optimizing language models for dialogue,” 2022,
OpenAI, · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford et al., · 2023
Cited alongside, same era.
“Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization,”
Puyuan Peng et al., · 2023
Cited alongside, same era.
“Lyricwhiz: Robust multilingual zero-shot lyrics transcription by whispering to chatgpt,”
Le Zhuo et al., · 2023
Cited alongside, same era.
“Gpt-4 technical report,” 2023
OpenAI, · 2023
Cited alongside, same era.
“Brouhaha: multi-task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,”
Marvin Lavechin et al., · 2023
Cited alongside, same era.
“Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,”
Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, et al., · 2024
Closest in time.
“Planning, creation, usage: Benchmarking LLMs for comprehensive tool utilization in real-world complex scenarios,”
Shijue Huang et al., · 2024
Closest in time.
“StableToolBench: Towards stable large-scale benchmarking on tool learning of large language models,”
Zhicheng Guo et al., · 2024
Closest in time.
“EASYTOOL: Enhancing LLM-based agents with concise tool instruction,”
Siyu Yuan et al., · 2024
Closest in time.
Junjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang, Sixian Li, Guanyu Li, Xiaoran Fan, Qi Zhang, Tao Gui, and Xuanjing Huang, · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Commonaccent: Exploring large acoustic pretrained models for accent classification based on common voice,”
Juan Zuluaga-Gomez, Sara Ahmed, Danielius Visockas, and Cem Subakan, · 2023
Cited alongside, same era.
“Powerset multi-class cross entropy loss for neural speaker diarization,”
Alexis Plaquet and Hervé Bredin, · 2023
Cited alongside, same era.
“Understanding the planning of llm agents: A survey,”
Xu Huang et al., · 2024
Cited alongside, same era.
“Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies,”
Liangming Pan et al., · 2024
Cited alongside, same era.
“Selfcheck: Using LLMs to zero-shot check their own step-by-step reasoning,”
Ning Miao, Yee Whye Teh, and Tom Rainforth, · 2024
Cited alongside, same era.
“Ai-augmented predictions: Llm assistants improve human forecasting accuracy,”
Philipp Schoenegger, Peter S Park, Ezra Karger, and Philip E Tetlock, · 2024
Cited alongside, same era.
Closest in time.
“Learning to use tools via cooperative and interactive agents,”
Zhengliang Shi et al., · 2024
Closest in time.
“Toolllm: Facilitating large language models to master 16000+ real-world apis,”
Yujia Qin et al., · 2024
Closest in time.
“Toolformer: Language models can teach themselves to use tools,”
Timo Schick et al., · 2024
Closest in time.
“Salmonn: Towards generic hearing abilities for large language models,”
Changli Tang et al., · 2024
Closest in time.
“Wavllm: Towards robust and adaptive speech large language model,”
Shujie Hu et al., · 2024
Closest in time.
“Desta: Enhancing speech language models through descriptive speech-text alignment,”
Ke-Han Lu et al., · 2024
Closest in time.
“Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,”
Zhifeng Kong et al., · 2024
Closest in time.
“Investigating zero-shot generalizability on mandarin-english code-switched asr and speech-to-text translation of recent foundation models with self-supervision and weak supervision,”
Chih-Kai Yang, Kuan-Po Huang, Ke-Han Lu, Chun-Yi Kuan, Chi-Yuan Hsiao, and Hung-Yi Lee, · 2024
Closest in time.
“Do prompts really prompt? exploring the prompt understanding capability of whisper,”
Chih-Kai Yang, Kuan-Po Huang, and Hung-yi Lee, · 2024
Closest in time.
“emotion2vec: Self-supervised pre-training for speech emotion representation,”
Ziyang Ma et al., · 2024
Closest in time.
“PhonologyBench: Evaluating phonological skills of large language models,”
Ashima Suvarna, Harshita Khandelwal, and Nanyun Peng, · 2024
Closest in time.
“Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models,”
Chun-Yi Kuan, Wei-Ping Huang, and Hung-yi Lee, · 2024
Closest in time.
“Zero resource code-switched speech benchmark using speech utterance pairs for multiple spoken languages,”
Kuan-Po Huang, Chih-Kai Yang, Yu-Kuan Fu, Ewan Dunbar, and Hung-yi Lee, · 2024
Closest in time.