Fetching the paper…
Reading the bibliography…
Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience compared to traditional text-based interactions.
Distortion measures for speech processing
R. Gray, A. Buzo, A. Gray, and Y. Matsuyama. 1980 · 1980
Earlier work this paper cites.
Stages in sentence production: An analysis of speech error data
Gary S Dell and Peter A Reich. 1981 · 1981
Earlier work this paper cites.
Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator
Y. Ephraim and D. Malah. 1984 · 1984
Earlier work this paper cites.
Preliminaries to a theory of speech disfluencies
Elizabeth Ellen Shriberg. 1994 · 1994
Earlier work this paper cites.
Grammar and the spoken language
Ronald Carter and Michael Mncarthy. 1995 · 1995
Earlier work this paper cites.
Spoken grammar: what is it and how can we teach it?
Michael McCarthy and Ronald Carter. 1995 · 1995
Earlier work this paper cites.
The effects of false starts and repetitions on the processing of subsequent words in spontaneous speech
Jean E.Fox Tree. 1995 · 1995
Earlier work this paper cites.
Acoustic properties of naturally produced clear speech at normal speaking rates
Jean C Krause and Louis D Braida. 2004 · 2004
Earlier work this paper cites.
Perceptual adaptation to non-native speech
Ann R Bradlow and Tessa Bent. 2008 · 2008
Earlier work this paper cites.
Microphone array processing for distant speech recognition: From close-talking microphones to far-field sensors
Kenichi Kumatani, John McDonough, and Bhiksha Raj. 2012 · 2012
Earlier work this paper cites.
Making machines understand us in reverberant rooms: Robustness against reverberation for automatic speech recognition
Takuya Yoshioka, Armin Sehr, Marc Delcroix, Keisuke Kinoshita, Roland Maas, Tomohiro Nakatani, and Walter Kellermann. 2012 · 2012
Earlier work this paper cites.
Speech recognition in natural background noise
Julien Meyer, Laure Dentel, and Fanny Meunier. 2013 · 2013
Earlier work this paper cites.
Packet loss concealment based on deep neural networks for digital speech transmission
Bong-Ki Lee and Joon-Hyuk Chang. 2016 · 2016
Earlier work this paper cites.
Multichannel signal processing with deep neural networks for automatic speech recognition
Tara N. Sainath, Ron J. Weiss, Kevin W. Wilson, Bo Li, Arun Narayanan, Ehsan Variani, Michiel Bacchiani, Izhak Shafran, Andrew W. Senior, Kean K. Chin, Ananya Misra, and Chanwoo Kim. 2017 · 2017
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. 2020 · 2020
Earlier work this paper cites.
Grammatical error detection in transcriptions of spoken English
Andrew Caines, Christian Bentz, Kate Knill, Marek Rei, and Paula Buttery. 2020 · 2020
Earlier work this paper cites.
TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020 · 2020
Cited alongside, same era.
SD-QA: Spoken dialectal question answering for the real world
Fahim Faisal, Sharlina Keshava, Md Mahfuz Ibn Alam, and Antonios Anastasopoulos. 2021 · 2021
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. 2023 · 2023
Cited alongside, same era.
Automatic pronunciation assessment - a review
Moshi: a speech-text foundation model for real-time dialogue
Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024 · 2024
Closest in time.
Zhihao Du, Qian Chen, Shiliang Zhang, Kai Hu, Heng Lu, Yexin Yang, Hangrui Hu, Siqi Zheng, Yue Gu, Ziyang Ma, et al. 2024 · 2024
Closest in time.
Llama-omni: Seamless speech interaction with large language models
Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma, Shaolei Zhang, and Yang Feng. 2024 · 2024
Closest in time.
Vita: Towards open-source interactive omni multimodal llm
Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Xiong Wang, Di Yin, Long Ma, Xiawu Zheng, et al. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yassine Kheir, Ahmed Ali, and Shammur Chowdhury. 2023 · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Cited alongside, same era.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023 · 2023
Cited alongside, same era.
Disfluency generation for more robust dialogue systems
Benjamin Marie. 2023 · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Cited alongside, same era.
Large language models can be easily distracted by irrelevant context
Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. 2023 · 2023
Cited alongside, same era.
SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023 · 2023
Cited alongside, same era.
Listen, think, and understand
Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, and James R. Glass. 2024 · 2024
Closest in time.
Distilling an end-to-end voice assistant without instruction training data
William Held, Ella Li, Michael Ryan, Weiyan Shi, Yanzhe Zhang, and Diyi Yang. 2024 · 2024
Closest in time.
Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech
Chien-yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Haibin Wu, Siddhant Arora, Kai-Wei Chang, Jiatong Shi, Yifan Peng, et al. 2024 · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2024 · 2024
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024 · 2024
Closest in time.
Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al. 2024 · 2024
Closest in time.
SafeDecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. 2024 · 2024
Closest in time.
Air-bench: Benchmarking large audio-language models via generative comprehension
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, et al. 2024 · 2024
Closest in time.
Evaluating large language models at evaluating instruction following
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2024 · 2024
Closest in time.
A chat about boring problems: Studying gpt-based text normalization
Yang Zhang, Travis M Bartley, Mariana Graterol-Fuenmayor, Vitaly Lavrukhin, Evelina Bakhturina, and Boris Ginsburg. 2024 · 2024
Closest in time.
End-to-end speech recognition and disfluency removal
Paria Jamshid Lou and Mark Johnson. 2020 · 2061
Closest in time.