Linguistic unit discovery from multi-modal inputs in unwritten languages: Summary of the ”Speaking Rosetta” JSALT 2017 workshop
Odette Scharenborg, Laurent Besacier, Alan W. Black, Mark Hasegawa-Johnson, Florian Metze, Graham Neubig, Sebastian Stüker, Pierre Godard, Markus Müller, Lucas Ondel, Shruti Palaskar, Philip Arthur, Francesco Ciannella, Mingxing Du, Elin Larsen, Danny Merkx, Rachid Riad, Liming Wang, and Emmanuel Dupoux · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Original
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Towards bilingual lexicon discovery from visually grounded speech audio
Emmanuel Azuh, David Harwath, and James Glass · 2019
Closest in time.
VQVAE with speaker adversarial training, 2019
Suhee Cho, Yeonjung Hong, Yookyunk Shin, and Youngsun Cho · 2019
Closest in time.
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron J. Weiss, Samy Bengio, and Aäron van den Oord · 2019
Closest in time.
Symbolic inductive bias for visually grounded learning of spoken language
Grzegorz Chrupała · 2019
Closest in time.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James R. Glass · 2019
Closest in time.
The zero resource speech challenge 2019: TTS without T
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel, Alan W. Black, Laurent Besacier, Sakriani Sakti, and Emmanuel Dupoux · 2019
Closest in time.
Combining adversarial training and disentangled speech representation for robust zero-resource subword modeling
Siyuan Feng, Tan Lee, and Zhiyuan Peng · 2019
Closest in time.
Google cloud speech-to-text API
Google · 2019
Closest in time.
Towards visually grounded sub-word speech unit discovery
David Harwath and James Glass · 2019
Closest in time.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adrià Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass · 2019
Closest in time.
Learning from multiview correlations in open-domain videos
Nils Holzenberger, Shruti Palaskar, Pranava Madhyastha, Florian Metze, and Raman Arora · 2019
Closest in time.
Transfer learning from audio-visual grounding to speech recognition
Wei-Ning Hsu, David Harwath, and James Glass · 2019
Closest in time.
Large-scale representation learning from visually grounded untranscribed speech
Gabriel Ilharco, Yuan Zhang, and Jason Baldridge · 2019
Closest in time.
Unsupervised end-to-end learning of discrete linguistic units for voice conversion
Andy T Liu, Po-chun Hsu, and Hung-yi Lee · 2019
Closest in time.
Language learning using speech to image retrieval
Danny Merkx, Stefan L. Frank, and Mirjam Ernestus · 2019
Closest in time.
On the contributions of visual and textual supervision in low-resource semantic speech retrieval
Ankita Pasad, Bowen Shi, Herman Kamper, and Karen Livescu · 2019
Closest in time.
Learning problem-agnostic speech representations from multiple self-supervised tasks
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio · 2019
Closest in time.
Generating diverse high-fidelity images with vq-vae-2
Original
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 2019
Closest in time.
Learning words by drawing images
Dídac Surís, Adrià Recasens, David Bau, David Harwath, James Glass, and Antonio Torralba · 2019
Closest in time.