Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelling techniques to audio data.
J. MacQueen, “Some methods for classification and analysis of multivariate observations,” in Berkeley Symposium on Mathematical Statistics and Probability , vol. 1, no. 14, 1967, pp. 281–297
1967
Earlier work this paper cites.
Y. Shibata, T. Kida, S. Fukamachi, M. Takeda, A. Shinohara, T. Shinohara, and S. Arikawa, “Byte Pair Encoding: A text compression scheme that accelerates pattern matching,” Technical Report, Department of Informatics, Kyushu University , 1999
1999
Earlier work this paper cites.
K. C. Pohlmann, Principles of Digital Audio . McGraw-Hill Professional, 2000
2000
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The million song dataset,” International Society for Music Information Retrieval Conference , pp. 591–596, 2011
2011
Earlier work this paper cites.
J. Sterne, MP3: The Meaning of a Format . Duke University Press, 2012
2012
Earlier work this paper cites.
J. M. Valin, K. Vos, and T. B. Terriberry, “Definition of the opus audio codec,” Internet Engineering Task Force Standard , 2012
2012
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. Terriberry, “Definition of the opus audio codec,” Tech. Rep., 2012
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” arXiv preprint:1312.6114 , 2013
2013
Earlier work this paper cites.
T. Brychcín and M. Konopík, “Semantic spaces for improving language modeling,” Computer Speech & Language , vol. 28, no. 1, pp. 192–209, 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in Neural Information Processing Systems , vol. 27, 2014
2014
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition,” arXiv preprint:1402.1128 , 2014
2014
Earlier work this paper cites.
R. M. Bittner, J. Salamon, M. Tierney, M. Mauch, C. Cannam, and J. P. Bello, “Medleydb: A multitrack dataset for annotation-intensive mir research.” in ISMIR , vol. 14, 2014, pp. 155–160
2014
Earlier work this paper cites.
ITU-R Recommendation BS.1534-3, “Method for the subjective assessment of intermediate quality level of audio systems,” International Telecommunication Union , 2014. [Online]. Available: https://www.itu.int/dms_pubrec/itu-r/rec/bs/R-REC-BS.1534-3-201510-I!!PDF-E.pdf
2014
Earlier work this paper cites.
H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE Transactions on Affective Computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Earlier work this paper cites.
M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache et al. , “Overview of the evs codec architecture,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2015, pp. 5698–5702
2015
Earlier work this paper cites.
D. Rezende and S. Mohamed, “Variational inference with normalizing flows,” in International Conference on Machine Learning . Proceedings of Machine Learning Research, 2015, pp. 1530–1538
2015
Earlier work this paper cites.
K. J. Piczak, “ESC: Dataset for environmental sound classification,” in Proceedings of the ACM International Conference on Multimedia , 2015, pp. 1015–1018
2015
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “ViSQOL: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, pp. 1–18, 2015
2015
Earlier work this paper cites.
A. O. Bayer and G. Riccardi, “Semantic language models with deep neural networks,” Computer Speech & Language , vol. 40, pp. 1–22, 2016
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” Dec. 2017. [Online]. Available: https://doi.org/10.5281/zenodo.1117372
2017
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “AudioSet: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2017, pp. 776–780
2017
Earlier work this paper cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with wavenet autoencoders,” in Proceedings of International Conference on Machine Learning , 2017, pp. 1068–1077
2017
Earlier work this paper cites.
A. Creswell, T. White, V. Dumoulin, K. Arulkumaran, B. Sengupta, and A. A. Bharath, “Generative adversarial networks: An overview,” IEEE Signal Processing Magazine , vol. 35, no. 1, pp. 53–65, 2018
2018
Earlier work this paper cites.
F.-R. Stöter, S. Chakrabarty, E. Habets, and B. Edler, “LibriCount: A dataset for speaker count estimation,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2018
2018
Earlier work this paper cites.
B. Kim, M. Ghei, B. Pardo, and Z. Duan, “Vocal Imitation Set: a dataset of vocally imitated sound events using the audioset ontology.” in Workshop on Detection and Classification of Acoustic Scenes and Events , 2018, pp. 148–152
2018
Earlier work this paper cites.
P. Warden, “Speech Commands: A dataset for limited-vocabulary speech recognition,” arXiv preprint:1804.03209 , 2018
2018
Cited alongside, same era.
C. Gârbacea, A. van den Oord, Y. Li, F. S. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “Low bit-rate speech coding with vq-vae and a wavenet decoder,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2019, pp. 735–739
2019
Cited alongside, same era.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “LibriTTS: A corpus derived from librispeech for text-to-speech,” in INTERSPEECH , 2019, pp. 1526–1530
2019
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems , vol. 33, pp. 6840–6851, 2020
2020
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 684–10 695
2022
Later among the works it cites.
A. Q. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. Mcgrew, I. Sutskever, and M. Chen, “GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models,” in International Conference on Machine Learning , 2022, pp. 16 784–16 804
2022
Later among the works it cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research , 2023
2023
Later among the works it cites.
D. Yang, S. Liu, R. Huang, J. Tian, C. Weng, and Y. Zou, “HiFi-Codec: Group-residual vector quantization for high fidelity audio codec,” arXiv preprint:2305.02765 , 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 6840–6851
2020
Cited alongside, same era.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, pp. 17 022–17 033, 2020
2020
Cited alongside, same era.
H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A large-scale audio-visual dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2020, pp. 721–725
2020
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
S. Sun, K. Krishna, A. Mattarella-Micke, and M. Iyyer, “Do long-range language models actually use long-range context?” in Conference on Empirical Methods in Natural Language Processing , 2021, pp. 807–822
2021
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi et al. , “AudioLM: A language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 42, pp. 2523–2544, 2023
2023
Later among the works it cites.
C. Toraman, E. H. Yilmaz, F. Şahinuç, and O. Ozcelik, “Impact of tokenization on language models: An analysis for turkish,” ACM Transactions on Asian and Low-Resource Language Information Processing , vol. 22, no. 4, pp. 1–21, 2023
2023
Later among the works it cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv preprint:2302.13971 , 2023
2023
Later among the works it cites.
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi et al. , “MusicLM: Generating music from text,” arXiv preprint:2301.11325 , 2023
2023
Later among the works it cites.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Defossez, “Simple and controllable music generation,” in Advances in Neural Information Processing Systems , vol. 36, 2023, pp. 47 704–47 720
2023
Later among the works it cites.
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al. , “Neural codec language models are zero-shot text to speech synthesizers,” arXiv preprint:2301.02111 , 2023
2023
Later among the works it cites.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in International Conference on Learning Representations , 2023
2023
Later among the works it cites.
P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. d. C. Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov et al. , “AudioPaLM: A large language model that can speak and listen,” arXiv preprint:2306.12925 , 2023
2023
Later among the works it cites.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to-audio generation with latent diffusion models,” International Conference on Machine Learning , 2023
2023
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: audio pre-training with acoustic tokenizers,” in Proceedings of International Conference on Machine Learning , 2023, pp. 5178–5193
2023
Later among the works it cites.
R. Sheffer and Y. Adi, “I hear your true colors: Image guided audio generation,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2023
2023
Later among the works it cites.
D. Ghosal, N. Majumder, A. Mehrish, and S. Poria, “Text-to-audio generation using instruction guided latent diffusion model,” in ACM International Conference on Multimedia , 2023, p. 3590–3598
2023
Later among the works it cites.
R. Huang, J. Huang, D. Yang, Y. Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-An-Audio: Text-to-audio generation with prompt-enhanced diffusion models,” International Conference on Machine Learning , 2023
2023
Later among the works it cites.
K. Zaman, M. Sah, C. Direkoglu, and M. Unoki, “A survey of audio classification using deep learning,” IEEE Access , vol. 11, pp. 106 620–106 649, 2023
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proceedings of International Conference on Machine Learning , 2023, pp. 28 492–28 518
2023
Later among the works it cites.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
X. Tan, J. Chen, H. Liu, J. Cong, C. Zhang, Y. Liu, X. Wang, Y. Leng, Y. Yi, L. He, S. Zhao, T. Qin, F. Soong, and T.-Y. Liu, “NaturalSpeech: End-to-end text-to-speech synthesis with human-level quality,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 6, pp. 4234–4245, 2024
2024
Closest in time.
K. Chen, Y. Wu, H. Liu, M. Nezhurina, T. Berg-Kirkpatrick, and S. Dubnov, “MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2024, pp. 1206–1210
2024
Closest in time.
H. Liu, K. Chen, Q. Tian, W. Wang, and M. D. Plumbley, “AudioSR: Versatile audio super-resolution at scale,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 2024, pp. 1076–1080
2024
Closest in time.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 2871–2883, 2024
2024
Closest in time.
S. Lin, B. Liu, J. Li, and X. Yang, “Common diffusion noise schedules and sample steps are flawed,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 5404–5411
2024
Closest in time.
X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang, “WavCaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 32, pp. 3339–3354, 2024
2024
Closest in time.