Fetching the paper…
Reading the bibliography…
We introduce AudioLM, a framework for high-quality audio generation with long-term consistency.
T. Schatz, V. Peddinti, F. R. Bach, A. Jansen, H. Hermansky, and E. Dupoux, “Evaluating speech features with the minimal-pair ABX task: analysis of the classical MFC/PLP pipeline,” in Interspeech . ISCA, 2013, pp. 1781–1785
2013
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram, and N. Harte, “ViSQOL: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 1, pp. 1–18, 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
T. Schatz, “Abx-discriminability measures and applications. (mesures de discriminabilité ABX et applications),” Ph.D. dissertation, Pierre and Marie Curie University, Paris, France, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017, pp. 6306–6315
2017
Earlier work this paper cites.
E. Agustsson, F. Mentzer, M. Tschannen, L. Cavigelli, R. Timofte, L. Benini, and L. V. Gool, “Soft-to-hard vector quantization for end-to-end learned compression of images and neural networks,” Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in International Conference on Machine Learning (ICML) , vol. 80. PMLR, 2018, pp. 2415–2424
2018
Earlier work this paper cites.
A. van den Oord, Y. Li, I. Babuschkin, K. Simonyan, O. Vinyals, K. Kavukcuoglu, G. van den Driessche, E. Lockhart, L. C. Cobo, F. Stimberg, N. Casagrande, D. Grewe, S. Noury, S. Dieleman, E. Elsen, N. Kalchbrenner, H. Zen, A. Graves, H. King, T. Walters, D. Belov, and D. Hassabis, “Parallel WaveNet: Fast high-fidelity speech synthesis,” in International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 80. PMLR, 2018, pp. 3915–3923
2018
Earlier work this paper cites.
S. Kankanahalli, “End-to-end optimized speech coding with deep neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 2521–2525
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brebisson, Y. Bengio, and A. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT) . Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Razavi, A. van den Oord, and O. Vinyals, “Generating diverse high-fidelity images with VQ-VAE-2,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019, pp. 14 837–14 847
2019
Earlier work this paper cites.
K. Zhen, J. Sung, M. S. Lee, S. Beack, and M. Kim, “Cascaded cross-module residual learning towards lightweight end-to-end speech coding,” in Interspeech . ISCA, 2019, pp. 3396–3400
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C. A. Huang, S. Dieleman, E. Elsen, J. H. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations (ICLR) , 2019
2019
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Earlier work this paper cites.
M. Tagliasacchi, Y. Li, K. Misiunas, and D. Roblek, “SEANet: A multi-modal speech enhancement network,” in Interspeech , 2020
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Advances in Neural Information Processing Systems (NeurIPS) , 2020
2020
Earlier work this paper cites.
M. Binkowski, J. Donahue, S. Dieleman, A. Clark, E. Elsen, N. Casagrande, L. C. Cobo, and K. Simonyan, “High fidelity speech synthesis with adversarial networks,” in International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
G. Lample and F. Charton, “Deep learning for symbolic mathematics,” in International Conference on Learning Representations (ICLR) , 2020
2020
Cited alongside, same era.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Trans. Assoc. Comput. Linguistics , vol. 9, pp. 53–68, 2021
2021
Later among the works it cites.
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlós, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller, “Rethinking attention with performers,” in International Conference on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . Computer Vision Foundation / IEEE, 2021, pp. 12 873–12 883
2021
Later among the works it cites.
M. Rivière and E. Dupoux, “Towards unsupervised learning of speech features in the wild,” in IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 156–163
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” JMLR , vol. 21, pp. 140:1–140:67, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Interspeech . ISCA, 2020, pp. 5036–5040
2020
Cited alongside, same era.
M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: an open source production ready objective speech and audio metric,” in Twelfth International Conference on Quality of Multimedia Experience (QoMEX) , 2020, pp. 1–6
2020
Cited alongside, same era.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-light: A benchmark for ASR with limited or no supervision,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7669–7673
2020
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Interspeech . ISCA, 2020, pp. 5036–5040
2020
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” in International Conference on Learning Representations (ICLR) , 2021
2021
Cited alongside, same era.
W. Hsu, Y. H. Tsai, B. Bolte, R. Salakhutdinov, and A. Mohamed, “Hubert: How much can a bad teacher benefit ASR pre-training?” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6533–6537
2021
Later among the works it cites.
B. van Niekerk, L. Nortje, M. Baas, and H. Kamper, “Analyzing speaker information in self-supervised models to improve zero-resource speech processing,” in Interspeech . ISCA, 2021, pp. 1554–1558
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
J. Yu, X. Li, J. Y. Koh, H. Zhang, R. Pang, J. Qin, A. Ku, Y. Xu, J. Baldridge, and Y. Wu, “Vector-quantized image modeling with improved VQGAN,” in International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, B. Hutchinson, W. Han, Z. Parekh, X. Li, H. Zhang, J. Baldridge, and Y. Wu, “Scaling autoregressive models for content-rich text-to-image generation,” Transactions on Machine Learning Research (TMLR) , 2022
2022
Closest in time.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE ACM Trans. Audio Speech Lang. Process. , vol. 30, pp. 495–507, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, M. M. Botvinick, I. Simon, H. Sheahan, N. Zeghidour, J. Alayrac, J. Carreira, and J. H. Engel, “General-purpose, long-context autoregressive modeling with Perceiver AR,” in International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 162. PMLR, 2022, pp. 8535–8558
2022
Closest in time.
2022
Closest in time.
E. Kharitonov, A. Lee, A. Polyak, Y. Adi, J. Copet, K. Lakhotia, T. A. Nguyen, M. Rivière, A. Mohamed, E. Dupoux, and W. Hsu, “Text-free prosody-aware generative spoken language modeling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL) . Association for Computational Linguistics, 2022, pp. 8666–8681
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
E. Casanova, J. Weber, C. D. Shulby, A. C. Júnior, E. Gölge, and M. A. Ponti, “Yourtts: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,” in International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 162. PMLR, 2022, pp. 2709–2720
2022
Closest in time.