Fetching the paper…
Reading the bibliography…
Inspired by the impressive capabilities of GPT-4o, there is growing interest in enabling speech language models (SLMs) to engage in natural, fluid spoken interactions with humans.
Fisher english training speech part 1 transcripts
Cieri, C., Graff, D., Kimball, O., Miller, D., and Walker, K · 2004
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Heafield, K · 2011
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and kavukcuoglu, k · 2017
Earlier work this paper cites.
Graphrnn: Generating realistic graphs with deep auto-regressive models
You, J., Ying, R., Ren, X., Hamilton, W. L., and Leskovec, J · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2019
Earlier work this paper cites.
Molecular transformer: A model for uncertainty-calibrated chemical reaction prediction
Schwaller, P., Laino, T., Gaudin, T., Bolgar, P., Hunter, C. A., Bekas, C., and Lee, A. A · 2019
Earlier work this paper cites.
Wav2vec 2.0: Learning the structure of speech from raw audio
Baevski, A., Auli, M., and Conneau, A · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
Graphaf: a flow-based autoregressive model for molecular graph generation
Shi, C., Xu, M., Zhu, Z., Zhang, W., Zhang, M., and Tang, J · 2020
Earlier work this paper cites.
Scaling autoregressive video models
Weissenborn, D., Täckström, O., and Uszkoreit, J · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2021
Earlier work this paper cites.
Diffwave: A versatile diffusion model for audio synthesis
Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J. L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Earlier work this paper cites.
Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Ao, J., Wang, R., Zhou, L., Wang, C., Ren, S., Wu, Y., Liu, S., Ko, T., Li, Q., Zhang, Y., Wei, Z., Qian, Y., Li, J., and Wei, F · 2022
Earlier work this paper cites.
Yourtts: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone
Casanova, E., Weber, J., Shulby, C. D., Júnior, A. C., Gölge, E., and Ponti, M. A · 2022
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W · 2022
Earlier work this paper cites.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S. C. H · 2022
Earlier work this paper cites.
Evolutionary-scale prediction of atomic level protein structure with a language model
Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S., and Rives, A · 2022
Earlier work this paper cites.
Speechnet: Weakly supervised, end-to-end speech recognition at industrial scale
Tang, R., Kumar, K., Yang, G., Pandey, A., Mao, Y., Belyaev, V., Emmadi, M., Murray, G. C., Ture, F., and Lin, J · 2022
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2022
Earlier work this paper cites.
Arora, S., Futami, H., Jung, J., Peng, Y., Sharma, R. S., Kashiwagi, Y., Tsunoo, E., and Watanabe, S · 2023
Earlier work this paper cites.
Audiolm: A language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2023
Earlier work this paper cites.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2023
Earlier work this paper cites.
Pengi: An audio language model for audio tasks
Deshmukh, S., Elizalde, B., Singh, R., and Wang, H · 2023
Cited alongside, same era.
CLAP learning audio concepts from natural language supervision
Elizalde, B., Deshmukh, S., Ismail, M. A., and Wang, H · 2023
Cited alongside, same era.
Funasr: A fundamental end-to-end speech recognition toolkit
Gao, Z., Li, Z., Wang, J., Luo, H., Shi, X., Chen, M., Li, Y., Zuo, L., Du, Z., and Zhang, S · 2023
Cited alongside, same era.
Textually pretrained speech language models
Hassid, M., Remez, T., Nguyen, T. A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Défossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y · 2023
Cited alongside, same era.
Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Huang, R., Huang, J., Yang, D., Ren, Y., Liu, L., Li, M., Ye, Z., Liu, J., Yin, X., and Zhao, Z · 2023
Cited alongside, same era.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition
Zhou, X., Wang, J., Cui, Z., Zhang, S., Yan, Z., Zhou, J., and Zhou, C · 2023
Later among the works it cites.
URL https://chattts.com/
Chattts, 2024 · 2024
Later among the works it cites.
Seed-tts: A family of high-quality versatile speech generation models
Anastassiou, P., Chen, J., Chen, J., Chen, Y., Chen, Z., Chen, Z., Cong, J., Deng, L., Ding, C., Gao, L., Gong, M., Huang, P., Huang, Q., Huang, Z., Huo, Y., Jia, D., Li, C., Li, F., Li, H., Li, J., Li, X., Li, X., Liu, L., Liu, S., Liu, S., Liu, X., Liu, Y., Liu, Z., Lu, L., Pan, J., Wang, X., Wang, Y., Wang, Y., Wei, Z., Wu, J., Yao, C., Yang, Y., Yi, Y., Zhang, J., Zhang, Q., Zhang, S., Zhang, W., Zhang, Y., Zhao, Z., Zhong, D., and Zhuang, X · 2024
Later among the works it cites.
VALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers
Chen, S., Liu, S., Zhou, L., Liu, Y., Tan, X., Li, J., Zhao, S., Qian, Y., and Wei, F · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Cited alongside, same era.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2023
Cited alongside, same era.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., and Hsu, W · 2023
Cited alongside, same era.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S. C. H · 2023
Cited alongside, same era.
Audioldm: Text-to-audio generation with latent diffusion models
Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D. P., Wang, W., and Plumbley, M. D · 2023
Cited alongside, same era.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
Large language models generate functional protein sequences across diverse families
Madani, A., Krause, B., Greene, E. R., Subramanian, S., Mohr, B. P., Holton, J. M., Olmos, J. L., Xiong, C., Sun, Z. Z., Socher, R., Fraser, J. S., and Naik, N. V · 2023
Cited alongside, same era.
Chu, Y., Xu, J., Yang, Q., Wei, H., Wei, X., Guo, Z., Leng, Y., Lv, Y., He, J., Lin, J., Zhou, C., and Zhou, J · 2024
Later among the works it cites.
Moshi: a speech-text foundation model for real-time dialogue
Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., and Zeghidour, N · 2024
Later among the works it cites.
The llama 3 herd of models
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Rozière, B., Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., McConnell, C., Keller, C., Touret, C., Wu, C., Wong, C., Ferrer, C. C., Nikolaidis, C., Allonsius, D., Song, D., Pintz, D., Livshits, D., Esiobu, D., Choudhary, D., Mahajan, D., Garcia-Olano, D., Perino, D., Hupkes, D., Lakomkin, E., AlBadawy, E., Lobanova, E., Dinan, E., Smith, E. M., Radenovic, F., Zhang, F., Synnaeve, G., Lee, G., Anderson, G. L., Nail, G., Mialon, G., Pang, G., Cucurell, G., Nguyen, H., Korevaar, H., Xu, H., Touvron, H., Zarov, I., Ibarra, I. A., Kloumann, I. M., Misra, I., Evtimov, I., Copet, J., Lee, J., Geffert, J., Vranes, J., Park, J., Mahadeokar, J., Shah, J., van der Linde, J., Billock, J., Hong, J., Lee, J., Fu, J., Chi, J., Huang, J., Liu, J., Wang, J., Yu, J., Bitton, J., Spisak, J., Park, J., Rocca, J., Johnstun, J., Saxe, J., Jia, J., Alwala, K. V., Upasani, K., Plawiak, K., Li, K., Heafield, K., Stone, K., and et al · 2024
Later among the works it cites.
Llama-omni: Seamless speech interaction with large language models
Fang, Q., Guo, S., Zhou, Y., Ma, Z., Zhang, S., and Feng, Y · 2024
Later among the works it cites.
Audiochatllama: Towards general-purpose speech abilities for llms
Fathullah, Y., Wu, C., Lakomkin, E., Li, K., Jia, J., Shangguan, Y., Mahadeokar, J., Kalinli, O., Fuegen, C., and Seltzer, M · 2024
Later among the works it cites.
Vita: Towards open-source interactive omni multimodal llm, 2024
Fu, C., Lin, H., Long, Z., Shen, Y., Zhao, M., Zhang, Y., Dong, S., Wang, X., Yin, D., Ma, L., Zheng, X., He, R., Ji, R., Wu, Y., Shan, C., and Sun, X · 2024
Later among the works it cites.
Autoregressive image generation without vector quantization
Li, T., Tian, Y., Li, H., Deng, M., and He, K · 2024
Later among the works it cites.
Language model can listen while speaking
Ma, Z., Song, Y., Du, C., Cong, J., Chen, Z., Wang, Y., Wang, Y., and Chen, X · 2024
Later among the works it cites.
Spoken question answering and speech continuation using spectrogram-powered LLM
Nachmani, E., Levkovitch, A., Hirsch, R., Salazar, J., Asawaroengchai, C., Mariooryad, S., Rivlin, E., Skerry-Ryan, R. J., and Ramanovich, M. T · 2024
Later among the works it cites.
Spirit-lm: Interleaved spoken and written language model
Nguyen, T. A., Muller, B., Yu, B., Costa-jussà, M. R., Elbayad, M., Popuri, S., Duquenne, P., Algayres, R., Mavlyutov, R., Gat, I., Synnaeve, G., Pino, J., Sagot, B., and Dupoux, E · 2024
Later among the works it cites.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Shen, K., Ju, Z., Tan, X., Liu, E., Leng, Y., He, L., Qin, T., Zhao, S., and Bian, J · 2024
Later among the works it cites.
Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis
Siuzdak, H · 2024
Later among the works it cites.
SALMONN: towards generic hearing abilities for large language models
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models
Team, C · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Tian, K., Jiang, Y., Yuan, Z., Peng, B., and Wang, L · 2024
Later among the works it cites.
Beyond turn-based interfaces: Synchronous llms as full-duplex dialogue agents
Veluri, B., Peloquin, B. N., Yu, B., Gong, H., and Gollakota, S · 2024
Later among the works it cites.
Next-gpt: Any-to-any multimodal LLM
Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation, 2024
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Mini-omni: Language models can hear, talk while thinking in streaming, 2024
Xie, Z. and Wu, C · 2024
Later among the works it cites.
Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt
Yang, D., Liu, S., Huang, R., Weng, C., and Meng, H · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2024
Later among the works it cites.