Fetching the paper…
Reading the bibliography…
We introduce AudioPaLM, a large language model for speech understanding and generation.
JANUS-III: Speech-to-speech translation in multiple languages
A. Lavie, A. Waibel, L. Levin, M. Finke, D. Gates, M. Gavalda, T. Zeppenfeld, and P. Zhan · 1997
Earlier work this paper cites.
Verbmobil: Foundations of speech-to-speech translation
W. Wahlster · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
The ATR multilingual speech-to-speech translation system
S. Nakamura, K. Markov, H. Nakaiwa, G. Kikui, H. Kawai, T. Jitsuhiro, J.-S. Zhang, H. Yamamoto, E. Sumita, and S. Yamamoto · 2006
Earlier work this paper cites.
Covost 2: A massively multilingual speech-to-text translation corpus
C. Wang, A. Wu, and J. M. Pino · 2007
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2012
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
O. Bojar, C. Buck, C. Callison-Burch, C. Federmann, B. Haddow, P. Koehn, C. Monz, M. Post, R. Soricut, and L. Specia · 2013
Earlier work this paper cites.
Findings of the 2015 workshop on statistical machine translation
O. Bojar, R. Chatterjee, C. Federmann, B. Haddow, M. Huck, C. Hokamp, P. Koehn, V. Logacheva, C. Monz, M. Negri, M. Post, C. Scarton, L. Specia, and M. Turchi · 2015
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
R. Sennrich, B. Haddow, and A. Birch · 2016
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (WMT17)
O. Bojar, R. Chatterjee, C. Federmann, Y. Graham, B. Haddow, S. Huang, M. Huck, P. Koehn, Q. Liu, V. Logacheva, C. Monz, M. Negri, M. Post, R. Rubino, L. Specia, and M. Turchi · 2017
Earlier work this paper cites.
Low-resource speech recognition and keyword-spotting
M. J. Gales, K. M. Knill, and A. Ragni · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Findings of the 2018 conference on machine translation (WMT18)
O. Bojar, C. Federmann, M. Fishel, Y. Graham, B. Haddow, M. Huck, P. Koehn, and C. Monz · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
M. Post · 2018
Earlier work this paper cites.
Findings of the 2019 conference on machine translation (WMT19)
L. Barrault, O. Bojar, M. R. Costa-jussà, C. Federmann, M. Fishel, Y. Graham, B. Haddow, M. Huck, P. Koehn, S. Malmasi, C. Monz, M. Müller, S. Pal, M. Post, and M. Zampieri · 2019
Earlier work this paper cites.
Speech-to-speech translation between untranscribed unknown languages
A. Tjandra, S. Sakti, and S. Nakamura · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Earlier work this paper cites.
Findings of the 2020 conference on machine translation (WMT20)
L. Barrault, M. Biesialska, O. Bojar, M. R. Costa-jussà, C. Federmann, Y. Graham, R. Grundkiewicz, B. Haddow, M. Huck, E. Joanis, T. Kocmi, P. Koehn, C.-k. Lo, N. Ljubešić, C. Monz, M. Morishita, M. Nagata, T. Nakazawa, S. Pal, M. Post, and M. Zampieri · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
Large-scale adversarial training for vision-and-language representation learning
Z. Gan, Y.-C. Chen, L. Li, C. Zhu, Y. Cheng, and J. Liu · 2020
Cited alongside, same era.
Mls: A large-scale multilingual dataset for speech research
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
mslam: Massively multilingual joint pre-training for speech and text
A. Bapna, C. Cherry, Y. Zhang, Y. Jia, M. Johnson, Y. Cheng, S. Khanuja, J. Riesa, and A. Conneau · 2022
Later among the works it cites.
AudioLM: a language modeling approach to audio generation
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, and F. Wei · 2022
Later among the works it cites.
Self-supervised learning with random-projection quantizer for speech recognition
C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, and Y. Wu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
W2V-Bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu · 2021
Cited alongside, same era.
Violet: End-to-end video-language transformers with masked visual-token modeling
T.-J. Fu, L. Li, Z. Gan, K. Lin, W. Y. Wang, L. Wang, and Z. Liu · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Png bert: Augmented bert on phonemes and graphemes for neural tts
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu · 2021
Cited alongside, same era.
Transformer-based direct speech-to-speech translation with transcoder
T. Kano, S. Sakti, and S. Nakamura · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, et al · 2021
Cited alongside, same era.
Textless speech-to-speech translation on real data
A. Lee, H. Gong, P.-A. Duquenne, H. Schwenk, P.-J. Chen, C. Wang, S. Popuri, J. Pino, J. Gu, and W.-N. Hsu · 2021
Cited alongside, same era.
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Later among the works it cites.
Audiogen: Textually guided audio generation
F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi · 2022
Later among the works it cites.
Direct speech-to-speech translation with discrete units
A. Lee, P.-J. Chen, C. Wang, J. Gu, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W.-N. Hsu · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2022
Later among the works it cites.
Scaling up models and data with t5x
A. Roberts, H. W. Chung, A. Levskaya, G. Mishra, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, C. Hawthorne, A. Lewkowycz, A. Salcianu, M. van Zee, J. Austin, S. Goodman, L. B. Soares, H. Hu, S. Tsvyashchenko, A. Chowdhery, J. Bastings, J. Bulian, X. Garcia, J. Ni, A. Chen, K. Kenealy, J. H. Clark, S. Lee, D. Garrette, J. Lee-Thorp, C. Raffel, N. Shazeer, M. Ritter, M. Bosma, A. Passos, J. Maitin-Shepard, N. Fiedel, M. Omernick, B. Saeta, R. Sepassi, A. Spiridonov, J. Newlan, and A. Gesmundo · 2022
Later among the works it cites.
Learning audio-visual speech representation by masked multimodal cluster prediction, 2022
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed · 2022
Later among the works it cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang · 2022
Later among the works it cites.
Musiclm: Generating music from text
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank · 2023
Closest in time.
Palm 2 technical report
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. T. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. Clark, L. E. Shafey, Y. Huang, K. S. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. ’Abrego, J. Ahn, J. Austin, P. Barham, J. A. Botha, J. Bradbury, S. Brahma, K. M. Brooks, M. Catasta, Y. Cheng, C. Cherry, C. A. Choquette-Choo, A. Chowdhery, C. Crépy, S. Dave, M. Dehghani, S. Dev, J. Devlin, M. C. D’iaz, N. Du, E. Dyer, V. Feinberg, F. Feng, V. Fienber, M. Freitag, X. García, S. Gehrmann, L. González, G. Gur-Ari, S. Hand, H. Hashemi, L. Hou, J. Howland, A. R. Hu, J. Hui, J. Hurwitz, M. Isard, A. Ittycheriah, M. Jagielski, W. H. Jia, K. Kenealy, M. Krikun, S. Kudugunta, C. Lan, K. Lee, B. Lee, E. Li, M.-L. Li, W. Li, Y. Li, J. Li, H. Lim, H. Lin, Z.-Z. Liu, F. Liu, M. Maggioni, A. Mahendru, J. Maynez, V. Misra, M. Moussalem, Z. Nado, J. Nham, E. Ni, A. Nystrom, A. Parrish, M. Pellat, M. Polacek, A. Polozov, R. Pope, S. Qiao, E. Reif, B. Richter, P. Riley, A. Ros, A. Roy, B. Saeta, R. Samuel, R. M. Shelby, A. Slone, D. Smilkov, D. R. So, D. Sohn, S. Tokumine, D. Valter, V. Vasudevan, K. Vodrahalli, X. Wang, P. Wang, Z. Wang, T. Wang, J. Wieting, Y. Wu, K. Xu, Y. Xu, L. W. Xue, P. Yin, J. Yu, Q. Zhang, S. Zheng, C. Zheng, W. Zhou, D. Zhou, S. Petrov, and Y. Wu · 2023
Closest in time.
Soundstorm: Efficient parallel audio generation
Z. Borsos, M. Sharifi, D. Vincent, E. Kharitonov, N. Zeghidour, and M. Tagliasacchi · 2023
Closest in time.
Fleurs: Few-shot learning evaluation of universal representations of speech
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna · 2023
Closest in time.
Singsong: Generating musical accompaniments from singing
C. Donahue, A. Caillon, A. Roberts, E. Manilow, P. Esling, A. Agostinelli, M. Verzetti, I. Simon, O. Pietquin, N. Zeghidour, and J. H. Engel · 2023
Closest in time.
Textually pretrained speech language models
M. Hassid, T. Remez, T. A. Nguyen, I. Gat, A. Conneau, F. Kreuk, J. Copet, A. Défossez, G. Synnaeve, E. Dupoux, R. Schwartz, and Y. Adi · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
E. Kharitonov, D. Vincent, Z. Borsos, R. Marinier, S. Girgin, O. Pietquin, M. Sharifi, M. Tagliasacchi, and N. Zeghidour · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al · 2023
Closest in time.
When and why are pre-trained word embeddings useful for neural machine translation?
Y. Qi, D. Sachan, M. Felix, S. Padmanabhan, and G. Neubig · 2084
Closest in time.