Fetching the paper…
Reading the bibliography…
Efficient audio representations in a compressed continuous latent space are critical for generative audio modeling and Music Information Retrieval (MIR) tasks.
P. Charbonnier, L. Blanc-Feraud et al. , “Deterministic edge-preserving regularization in computed imaging,” IEEE Transactions on Image Processing , vol. 6, no. 2, pp. 298–311, 1997
1997
Earlier work this paper cites.
D. Wolff, S. Stober et al. , “A systematic comparison of music similarity adaptation approaches,” in Proceedings of the 13th International Society for Music Information Retrieval Conference, ISMIR 2012, Mosteiro S.Bento Da Vitória, Porto, Portugal, October 8-12, 2012 . FEUP Edições, 2012, pp. 103–108
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in 2nd International Conference on Learning Representations (ICLR) , Apr. 2014
2014
Earlier work this paper cites.
O. Ronneberger, P. Fischer et al. , “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 - 18th International Conference Munich, Germany, October 5 - 9, 2015, Proceedings, Part III , ser. Lecture Notes in Computer Science, vol. 9351. Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
A. Hines, J. Skoglund et al. , “Visqol: an objective speech quality model,” EURASIP J. Audio Speech Music. Process. , vol. 2015, p. 13, 2015
2015
Earlier work this paper cites.
A. van den Oord, O. Vinyals et al. , “Neural discrete representation learning,” in Advances in Neural Information Processing Systems 30 , Dec. 2017, pp. 6306–6315
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer et al. , “Attention is all you need,” in Advances in Neural Information Processing Systems 30 , Dec. 2017, pp. 5998–6008
2017
Earlier work this paper cites.
C. Sloan, N. Harte et al. , “Objective assessment of perceptual audio quality using visqolaudio,” IEEE Trans. Broadcast. , vol. 63, no. 4, pp. 693–705, 2017
2017
Earlier work this paper cites.
P. Ramachandran, B. Zoph et al. , “Searching for activation functions,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Workshop Track Proceedings . OpenReview.net, 2018
2018
Earlier work this paper cites.
Ángel Faraldo, “Beatport edm key dataset,” Jan. 2018
2018
Earlier work this paper cites.
A. Razavi, A. van den Oord et al. , “Generating diverse high-fidelity images with VQ-VAE-2,” in Advances in Neural Information Processing Systems 32 , Dec. 2019, pp. 14 837–14 847
2019
Earlier work this paper cites.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada , 2019, pp. 11 895–11 907
2019
Earlier work this paper cites.
Y. Song, S. Garg et al. , “Sliced Score Matching: A Scalable Approach to Density and Score Estimation,” in Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI 2019, Tel Aviv, Israel, July 22-25, 2019 , ser. Proceedings of Machine Learning Research, vol. 115. AUAI Press, 2019, pp. 574–584
2019
Earlier work this paper cites.
D. Bogdanov, M. Won et al. , “The mtg-jamendo dataset for automatic music tagging,” in Machine Learning for Music Discovery Workshop, International Conference on Machine Learning (ICML 2019) , Long Beach, CA, United States, 2019
2019
Earlier work this paper cites.
J. L. Roux, S. Wisdom et al. , “SDR - half-baked or well done?” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2019, Brighton, United Kingdom, May 12-17, 2019 . IEEE, 2019, pp. 626–630
2019
Earlier work this paper cites.
K. Kilgour, M. Zuluaga et al. , “Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms,” in 20th Annual Conference of the International Speech Communication Association (INTERSPEECH) , Sep. 2019, pp. 2350–2354
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Y. Song and S. Ermon, “Improved techniques for training score-based generative models,” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , 2020
2020
Cited alongside, same era.
J. Nistal, S. Lattner et al. , “DRUMGAN: synthesis of drum sounds with timbral feature conditioning using generative adversarial networks,” in Proceedings of the 21th International Society for Music Information Retrieval Conference (ISMIR) , Oct. 2020, pp. 590–597
2020
Cited alongside, same era.
——, “Comparing representations for audio synthesis using generative adversarial networks,” in 28th European Signal Processing Conference (EUSIPCO) . IEEE, Jan. 2020, pp. 161–165
2020
Cited alongside, same era.
L. Liu, H. Jiang et al. , “On the variance of the adaptive learning rate and beyond,” in 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net, 2020
Y. Song, P. Dhariwal et al. , “Consistency Models,” May 2023, arXiv:2303.01469 [cs, stat]
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
M. Chinen, F. S. C. Lim et al. , “Visqol v3: An open source production ready objective speech and audio metric,” in Twelfth International Conference on Quality of Multimedia Experience, QoMEX 2020, Athlone, Ireland, May 26-28, 2020 . IEEE, 2020, pp. 1–6
2020
Cited alongside, same era.
C. Emanuele, D. Ghisi et al. , “TinySOL: an audio dataset of isolated musical notes,” Jan. 2020
2020
Cited alongside, same era.
P. Esser, R. Rombach et al. , “Taming transformers for high-resolution image synthesis,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, Jun. 2021, pp. 12 873–12 883
2021
Cited alongside, same era.
J. Song, C. Meng et al. , “Denoising Diffusion Implicit Models,” in 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net, 2021
2021
Cited alongside, same era.
R. Castellon, C. Donahue et al. , “Codified audio language modeling learns useful representations for music information retrieval,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021, Online, November 7-12, 2021 , 2021, pp. 88–96
2021
Cited alongside, same era.
J. Spijkervet and J. A. Burgoyne, “Contrastive learning of musical representations,” in Proceedings of the 22nd International Society for Music Information Retrieval Conference, ISMIR 2021, Online, November 7-12, 2021 , 2021, pp. 673–681
2021
Cited alongside, same era.
M. Pasini and J. Schlüter, “Musika! Fast Infinite Waveform Music Generation,” in Proceedings of the 23rd International Society for Music Information Retrieval Conference, ISMIR 2022, Bengaluru, India, December 4-8, 2022 , 2022, pp. 543–550
2022
Cited alongside, same era.
N. Zeghidour, A. Luebs et al. , “SoundStream: An End-to-End Neural Audio Codec,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 30, pp. 495–507, 2022
2022
Cited alongside, same era.
Z. Ye, W. Xue et al. , “Comospeech: One-step speech and singing voice synthesis via consistency model,” in Proceedings of the 31st ACM International Conference on Multimedia, MM 2023, Ottawa, ON, Canada, 29 October 2023- 3 November 2023 . ACM, 2023, pp. 1831–1839
2023
Later among the works it cites.
J. Richter, S. Welker et al. , “Speech enhancement and dereverberation with diffusion-based generative models,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 31, pp. 2351–2364, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Wu, K. Chen et al. , “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023 . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Li, R. Yuan et al. , “Mert: Acoustic music understanding model with large-scale self-supervised training,” 2023
2023
Later among the works it cites.
C. Plachouras, “Beyond Benchmarks: A Toolkit for Music Audio Representation Evaluation,” Ph.D. dissertation, Universitat Pompeu Fabra, Sep. 2023
2023
Later among the works it cites.
C. Plachouras, P. Alonso-Jiménez et al. , “mir_ref: A representation evaluation framework for music information retrieval tasks,” in 37th Conference on Neural Information Processing Systems (NeurIPS), Machine Learning for Audio Workshop , New Orleans, LA, USA, 2023
2023
Later among the works it cites.
Z. Geng, W. Luo et al. , “Consistency models made easy,” 2024
2024
Closest in time.
2024
Closest in time.
M. Pasini, M. Grachten et al. , “Bass accompaniment generation via latent diffusion,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 1166–1170
2024
Closest in time.
2024
Closest in time.