Fetching the paper…
Reading the bibliography…
We present PAST, a novel end-to-end framework that jointly models phonetic information alongside signal reconstruction, eliminating the need for external pretrained models.
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, and V. Zue, “Timit acoustic-phonetic continuous speech corpus,” Linguistic Data Consortium , 11 1992
1992
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221) , vol. 2, 2001, pp. 749–752 vol.2
2001
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
T. Schatz et al. , “Evaluating speech features with the minimal-pair abx task: analysis of the classical mfc/plp pipeline,” in Interspeech 2013 , 2013, pp. 1781–1785
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in neural information processing systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
M. Chinen, F. S. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “Visqol v3: An open source production ready objective speech and audio metric,” in 2020 twelfth international conference on quality of multimedia experience (QoMEX) . IEEE, 2020, pp. 1–6
2020
Earlier work this paper cites.
T. A. Nguyen et al. , “The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” Self-Supervised Learning for Speech and Audio Processing Workshop @ NeurIPS , 2020
2020
Earlier work this paper cites.
K. Lakhotia et al. , “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1336–1354, 2021
2021
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM transactions on audio, speech, and language processing , vol. 29, pp. 3451–3460, 2021
2021
Earlier work this paper cites.
E. Kharitonov et al. , “Text-free prosody-aware generative spoken language modeling,” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pp. 8666–8681, 2022
2022
Earlier work this paper cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” arXiv preprint: 2210.13438 , 2022
2022
Cited alongside, same era.
S. Chen et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, p. 1505–1518, Oct. 2022
2022
Cited alongside, same era.
O. Tal, M. Mandel, F. Kreuk, and Y. Adi, “A systematic comparison of phonetic aware techniques for speech enhancement,” in Interspeech 2022 , 2022, pp. 1193–1197
2022
Cited alongside, same era.
R. Algayres et al. , “Generative spoken language model based on continuous word-sized audio tokens,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2023, pp. 3008–3028
2023
Cited alongside, same era.
M. Hassid et al. , “Textually pretrained speech language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
X. Zhang, D. Zhang, S. Li, Y. Zhou, and X. Qiu, “Speechtokenizer: Unified speech tokenizer for speech language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Later among the works it cites.
Z. Ye et al. , “Codec does matter: Exploring the semantic shortcoming of codec for audio language model,” arXiv preprint: 2408.17175 , 2024
2024
Later among the works it cites.
A. Défossez, L. Mazaré, M. Orsini, A. Royer, P. Pérez, H. Jégou, E. Grave, and N. Zeghidour, “Moshi: a speech-text foundation model for real-time dialogue,” Technical report, Kyutai , 2024
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Borsos et al. , “Audiolm: a language modeling approach to audio generation,” IEEE/ACM transactions on audio, speech, and language processing , vol. 31, pp. 2523–2533, 2023
2023
Cited alongside, same era.
A. Sicherman and Y. Adi, “Analysing discrete self supervised speech representation for spoken language modeling,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, p. 1–5
2023
Cited alongside, same era.
D. Yang, S. Liu, R. Huang, J. Tian, C. Weng, and Y. Zou, “Hifi-codec: Group-residual vector quantization for high fidelity audio codec,” arXiv preprint: 2305.02765 , 2023
2023
Cited alongside, same era.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research , 2023, featured Certification, Reproducibility Certification
2023
Cited alongside, same era.
Y.-C. Wu et al. , “Audiodec: An open-source streaming high-fidelity neural audio codec,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning . PMLR, 2023, pp. 28 492–28 518
2023
Cited alongside, same era.
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi, “Audiogen: Textually guided audio generation,” arXiv preprint: 2209.15352 , 2023
2023
Cited alongside, same era.
M.-J. Hwang et al. , “Textless acoustic model with self-supervised distillation for noise-robust expressive speech-to-speech translation,” Findings of the Association for Computational Linguistics: ACL 2024 , 2024
2024
Cited alongside, same era.
2024
Later among the works it cites.
Z. Du, S. Zhang, K. Hu, and S. Zheng, “Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 591–595
2024
Later among the works it cites.
S. Ji et al. , “Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling,” 2024
2024
Later among the works it cites.
A. Zeng, Z. Du, M. Liu, L. Zhang, S. Jiang, Y. Dong, and J. Tang, “Scaling speech-text pre-training with synthetic interleaved data,” arXiv preprint: 2411.17607 , 2024
2024
Later among the works it cites.
A. Turetzky and Y. Adi, “Last: Language model aware speech tokenization,” arXiv preprint: 2409.03701 , 2024
2024
Later among the works it cites.
S. Messica and Y. Adi, “Nast: Noise aware speech tokenization for speech language models,” arXiv preprint: 2406.11037 , 2024
2024
Later among the works it cites.
P. Mousavi, L. D. Libera, J. Duret, A. Ploujnikov, C. Subakan, and M. Ravanelli, “Dasb - discrete audio and speech benchmark,” arXiv preprint: 2406.14294 , 2024
2024
Later among the works it cites.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.