Fetching the paper…
Reading the bibliography…
We present SoundStorm, a model for efficient, non-autoregressive audio generation.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J., Bradbury, J., Xiong, C., Li, V. O., and Socher, R · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Ghazvininejad, M., Levy, O., Liu, Y., and Zettlemoyer, L · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, H., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R · 2020
Earlier work this paper cites.
Libri-Light: A benchmark for ASR with limited or no supervision
Kahn, J., Rivière, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., Likhomanenko, T., Synnaeve, G., Joulin, A., Mohamed, A., and Dupoux, E · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Earlier work this paper cites.
Rethinking attention with Performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlós, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A · 2021
Earlier work this paper cites.
w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Chung, Y., Zhang, Y., Han, W., Chiu, C., Qin, J., Pang, R., and Wu, Y · 2021
Earlier work this paper cites.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W., Bolte, B., Tsai, Y. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Earlier work this paper cites.
On generative spoken language modeling from raw audio
Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors
Reddy, C. K. A., Gopal, V., and Cutler, R · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Cited alongside, same era.
Nyströmformer: A Nyström-based algorithm for approximating self-attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Cited alongside, same era.
AudioGen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Later among the works it cites.
Autoregressive image generation using residual quantization
Lee, D., Kim, C., Kim, S., Cho, M., and Han, W.-S · 2022
Later among the works it cites.
Phenaki: Variable length video generation from open domain textual description
Villegas, R., Babaeizadeh, M., Kindermans, P.-J., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2022
Later among the works it cites.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L., Barua, A., Constant, N., Al-Rfou, R., Narang, S., Kale, M., Roberts, A., and Raffel, C · 2022
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AudioLM: A language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Cited alongside, same era.
MaskGIT: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Cited alongside, same era.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., Wu, J., Zhou, L., Ren, S., Qian, Y., Qian, Y., Wu, J., Zeng, M., Yu, X., and Wei, F · 2022
Cited alongside, same era.
Self-supervised learning with random-projection quantizer for speech recognition
Chiu, C., Qin, J., Zhang, Y., Yu, J., and Wu, Y · 2022
Cited alongside, same era.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
General-purpose, long-context autoregressive modeling with Perceiver AR
Hawthorne, C., Jaegle, A., Cangea, C., Borgeaud, S., Nash, C., Malinowski, M., Dieleman, S., Vinyals, O., Botvinick, M. M., Simon, I., Sheahan, H., Zeghidour, N., Alayrac, J., Carreira, J., and Engel, J. H · 2022
Cited alongside, same era.
MusicLM: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J. H., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., and Frank, C. H · 2023
Closest in time.
Singsong: Generating musical accompaniments from singing
Donahue, C., Caillon, A., Roberts, A., Manilow, E., Esling, P., Agostinelli, A., Verzetti, M., Simon, I., Pietquin, O., Zeghidour, N., and Engel, J. H · 2023
Closest in time.
Jouppi, N. P., Kurian, G., Li, S., Ma, P., Nagarajan, R., Nai, L., Patil, N., Subramanian, S., Swing, A., Towles, B., et al · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., and Wei, F · 2023
Closest in time.