Fetching the paper…
Reading the bibliography…
Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs).
D. A. Huffman, “A method for the construction of minimum-redundancy codes,” Proceedings of the IRE , vol. 40, no. 9, pp. 1098–1101, 1952
1952
Earlier work this paper cites.
R. Gray, “Vector quantization,” IEEE ASSP Magazine , vol. 1, no. 2, pp. 4–29, 1984
1984
Earlier work this paper cites.
J. S. Garofolo and et al., “Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,” NASA STI/Recon technical report n , vol. 93, p. 27403, 1993
1993
Earlier work this paper cites.
P. Gage, “A new algorithm for data compression,” C Users J. , vol. 12, no. 2, p. 23–38, feb 1994
1994
Earlier work this paper cites.
F. P. Mechel, “Acoustics of moving sources moving source,” 2008
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, E. A. Kazemzadeh, E. M. Provost, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, pp. 335–359, 2008. [Online]. Available: https://api.semanticscholar.org/CorpusID:11820063
2008
Earlier work this paper cites.
M. J. F. Gales and et al., “Speech recognition and keyword spotting for low-resource languages: Babel project research at cued,” in SLTU , 2014
2014
Earlier work this paper cites.
D. Renshaw, H. Kamper, A. Jansen, and S. Goldwater, “A comparison of neural network methods for unsupervised representation learning on the zero resource speech challenge,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Versteegh, X. Anguera, A. Jansen, and E. Dupoux, “The zero resource speech challenge 2015: Proposed approaches and results,” Procedia Computer Science , vol. 81, pp. 67–72, 2016
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , K. Erk and N. A. Smith, Eds. Berlin, Germany: Association for Computational Linguistics, Aug. 2016, pp. 1715–1725. [Online]. Available: https://aclanthology.org/P16-1162
2016
Earlier work this paper cites.
W.-N. Hsu, Y. Zhang, and J. Glass, “Unsupervised learning of disentangled and interpretable representations from sequential data,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” in Interspeech , 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:10475843
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised Pre-Training for Speech Recognition,” in Proc. Interspeech 2019 , 2019, pp. 3465–3469
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y.-A. Chung and J. Glass, “Generative pre-training for speech with autoregressive predictive coding,” in ICASSP 2020 . IEEE, 2020, pp. 3497–3501
2020
Cited alongside, same era.
S. Chen and et al., “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, 2022
2022
Later among the works it cites.
Z. Borsos and et al., “Audiolm: A language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2523–2533, 2022
2022
Later among the works it cites.
C.-C. Chiu and et al., “Self-supervised learning with random-projection quantizer for speech recognition,” in Proc. 39th ICML , ser. Proc. Mach. Learn. Res. PMLR, 2022, pp. 3915–3924
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proceedings of the 34th NeurIPS Conference , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” in Proc. Interspeech 2020 , 2020, pp. 2757–2761
2020
Cited alongside, same era.
2020
Cited alongside, same era.
K. W. Cheuk, Y.-J. Luo, E. Benetos, and D. Herremans, “The effect of spectrogram reconstruction on automatic music transcription: An alternative approach to improve transcription accuracy,” 2020 25th International Conference on Pattern Recognition (ICPR) , pp. 9091–9098, 2020. [Online]. Available: https://api.semanticscholar.org/CorpusID:224803213
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y.-A. Chung and et al., “w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” 2021 IEEE ASRU Workshop , pp. 244–250, 2021
2021
Cited alongside, same era.
A. Conneau and et al., “Fleurs: Few-shot learning evaluation of universal representations of speech,” 2022 IEEE Spoken Language Technology Workshop , pp. 798–805, 2022
2022
Later among the works it cites.
S. Dutta and S. Ganapathy, “Multimodal transformer with learnable frontend and self attention for emotion recognition,” ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 6917–6921, 2022. [Online]. Available: https://api.semanticscholar.org/CorpusID:249437569
2022
Later among the works it cites.
2023
Later among the works it cites.
R. A. et al., “Palm 2 technical report,” ArXiv , vol. abs/2305.10403, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Algayres and et al., “Generative spoken language model based on continuous word-sized audio tokens,” in EMNLP , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.