Fetching the paper…
Reading the bibliography…
We propose the SAMU-XLSR: Semantically-Aligned Multimodal Utterance-level Cross-Lingual Speech Representation learning framework.
G. Lample and A. Conneau, “Cross-lingual language model pretraining,” 2019. [Online]. Available:
1901
Earlier work this paper cites.
1904
Earlier work this paper cites.
1904
Earlier work this paper cites.
1907
Earlier work this paper cites.
1910
Earlier work this paper cites.
1911
Earlier work this paper cites.
1911
Earlier work this paper cites.
1911
Earlier work this paper cites.
2001
Earlier work this paper cites.
2002
Earlier work this paper cites.
2002
Earlier work this paper cites.
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
2007
Earlier work this paper cites.
2011
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in
2012
Cited alongside, same era.
A. Graves, “Sequence transduction with recurrent neural networks,”
2012
Cited alongside, same era.
2016
Cited alongside, same era.
D. Harwath, A. Torralba, and J. Glass, “Unsupervised learning of spoken language with visual context,”
2016
Cited alongside, same era.
2019
Later among the works it cites.
2020
Later among the works it cites.
Y.-A. Chung and J. Glass, “Generative pre-training for speech with autoregressive predictive coding,” in
2020
Later among the works it cites.
P. Safari, M. India, and J. Hernando, “Self-attention encoding and pooling for speaker recognition,”
2020
Later among the works it cites.
D. Harwath, W.-N. Hsu, and J. Glass, “Learning hierarchical discrete linguistic units from visually-grounded speech,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
M. A. Di Gangi, R. Cattoni, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: a Multilingual Speech Translation Corpus,” in
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in
2018
Cited alongside, same era.
2020
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao
2021
Later among the works it cites.
2021
Later among the works it cites.
P.-A. Duquenne, H. Gong, and H. Schwenk, “Multimodal and multilingual embeddings for large-scale speech mining,”
2021
Later among the works it cites.
2021
Later among the works it cites.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in
2021
Later among the works it cites.
S. Watanabe, F. Boyer, X. Chang, P. Guo, T. Hayashi, Y. Higuchi, T. Hori, W.-C. Huang, H. Inaguma, N. Kamo
2021
Later among the works it cites.
2022
Closest in time.
Wikipedia contributors, “Word error rate — Wikipedia, the free encyclopedia,” 2020, [Online; accessed 23-April-2022]. [Online]. Available:
2022
Closest in time.
S. Arora, S. Dalmia, P. Denisov, X. Chang, Y. Ueda, Y. Peng, Y. Zhang, S. Kumar, K. Ganesan, B. Yan
2022
Closest in time.