Fetching the paper…
Reading the bibliography…
Representing speech and audio signals in discrete units has become a compelling alternative to traditional high-dimensional feature vectors.
A. Graves et al. , “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
G. Hinton et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Process. Mag. , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
J. Chorowski et al. , “Attention-based models for speech recognition,” Proc. NeurIPS , vol. 2015, pp. 577–585, 2015
2015
Earlier work this paper cites.
T. N. Sainath et al. , “Learning the speech front-end with raw waveform CLDNNs,” Proc. Interspeech , 2015
2015
Earlier work this paper cites.
V. Panayotov et al. , “Librispeech: an asr corpus based on public domain audio books,” in Proc. IEEE ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
Y. Qian et al. , “Very deep convolutional neural networks for noise robust speech recognition,” IEEE/ACM Trans. ASLP. , vol. 24, no. 12, pp. 2263–2276, 2016
2016
Earlier work this paper cites.
K. Ito and L. Johnson, “The LJ speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
L. Dong et al. , “Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,” in Proc. IEEE ICASSP , 2018
2018
Earlier work this paper cites.
S. Watanabe et al. , “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech , 2018, pp. 2207–2211
2018
Earlier work this paper cites.
A. Gulati et al. , “Conformer: Convolution-augmented Transformer for speech recognition,” in Proc. Interspeech . ISCA, 2020, pp. 5036–5040
2020
Earlier work this paper cites.
A. Baevski et al. , “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Proc. NeurIPS , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Hayashi et al. , “ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,” in Proc. IEEE ICASSP . IEEE, 2020, pp. 7654–7658
2020
Earlier work this paper cites.
P. Lu et al. , “XiaoiceSing: A high-quality and integrated singing voice synthesis system,” Proc. Interspeech , 2020
2020
Earlier work this paper cites.
Y. Ren et al. , “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2020
2020
Earlier work this paper cites.
P. Guo et al. , “Recent developments on espnet toolkit boosted by conformer,” in Proc. IEEE ICASSP , 2021, pp. 5874–5878
2021
Earlier work this paper cites.
W.-N. Hsu et al. , “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. ASLP. , vol. 29, pp. 3451–3460, 2021
2021
Earlier work this paper cites.
K. Lakhotia et al. , “On generative spoken language modeling from raw audio,” Trans. ACL. , vol. 9, pp. 1336–1354, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
——, “ESPnet2-TTS: Extending the edge of tts research,” arXiv preprint arXiv:2110.07840 , 2021
2021
Cited alongside, same era.
A. Polyak et al. , “Speech Resynthesis from Discrete Disentangled Self-Supervised Representations,” in Proc. Interspeech , 2021, pp. 3615–3619
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
T. A. Nguyen et al. , “Generative spoken dialogue language modeling,” Trans. ACL. , vol. 11, pp. 250–266, 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
S. Chen et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE JSTSP , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
W. C. Huang et al. , “The VoiceMOS Challenge 2022,” in Proc. Interspeech , 2022, pp. 4536–4540
2022
Cited alongside, same era.
Y. Zhang et al. , “VISinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,” in Proc. IEEE ICASSP , 2022
2022
Cited alongside, same era.
Y. Wang et al. , “Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis,” in Proc. Interspeech , 2022, pp. 4242–4246
2022
Cited alongside, same era.
2023
Later among the works it cites.
J. Copet et al. , “Simple and controllable music generation,” arXiv preprint arXiv:2306.05284 , 2023
2023
Later among the works it cites.
W.-N. Hsu et al. , “Revise: Self-supervised speech resynthesis with visual input for universal and generalized speech regeneration,” in Proc. IEEE/CVF CVPR , 2023, pp. 18 795–18 805
2023
Later among the works it cites.
J. Shi et al. , “ML-SUPERB: Multilingual Speech Universal PERformance Benchmark,” in Proc. Interspeech , 2023, pp. 884–888
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
B. Yan et al. , “ESPnet-ST-v2: Multipurpose spoken language translation toolkit,” in Proc. ACL , Jul. 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Shi et al. , “Enhancing speech-to-speech translation with multiple TTS targets,” in Proc. IEEE ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
R. Kumar et al. , “High-fidelity audio compression with improved rvqgan,” Proc. NeurIPS , vol. 36, 2024
2024
Closest in time.