Fetching the paper…
Reading the bibliography…
In NLP, text language models based on words or subwords are known to outperform their character-based counterparts.
The zero resource speech challenge 2019: TTS without T
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel, Alan W. Black, Laurent Besacier, Sakriani Sakti, and Emmanuel Dupoux. 2019 · 1904
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli. 2019 · 1910
Earlier work this paper cites.
Libri-light: A benchmark for ASR with limited or no supervision
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, Tatiana Likhomanenko, Gabriel Synnaeve, Armand Joulin, Abdelrahman Mohamed, and Emmanuel Dupoux. 2019 · 1912
Earlier work this paper cites.
Class-based n
Peter F. Brown, Vincent J. Della Pietra, Peter V. deSouza, Jenifer C. Lai, and Robert L. Mercer. 1992 · 1992
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage. 1994 · 1994
Earlier work this paper cites.
Deep voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2017 · 2000
Earlier work this paper cites.
The buckeye corpus of conversational speech: labeling conventions and a test of transcriber reliability
Mark A. Pitt, Keith Johnson, Elizabeth Hume, Scott Kiesling, and William Raymond. 2005 · 2005
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2006
Earlier work this paper cites.
Chinese whispers - an efficient graph clustering algorithm and its application to natural language processing problems
Chris Biemann. 2006 · 2006
Earlier work this paper cites.
A review of vector quantization techniques
A. Vasuki and P.T. Vanathi. 2006 · 2006
Earlier work this paper cites.
Self-supervised contrastive learning for unsupervised phoneme segmentation
Felix Kreuk, Joseph Keshet, and Yossi Adi. 2020 · 2007
Earlier work this paper cites.
A bayesian framework for word segmentation: Exploring the effects of context
Sharon Goldwater, Thomas Griffiths, and Mark Johnson. 2009 · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2010
Earlier work this paper cites.
Revisiting k-means: New algorithms via bayesian nonparametrics
Brian Kulis and Michael I. Jordan. 2011 · 2011
Earlier work this paper cites.
Subword language modeling with neural networks
Tomas Mikolov, Ilya Sutskever, Anoop Deoras, Hai Son Le, Stefan Kombrink, and Jan Honzaernocky. 2011 · 2011
Earlier work this paper cites.
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Evgeny Kharitonov, Alexei Baevski, Ewan Dunbar, and Emmanuel Dupoux. 2020 · 2011
Earlier work this paper cites.
Crowdmos: An approach for crowdsourcing mean opinion score studies
Flávio Ribeiro, Dinei Florêncio, Cha Zhang, and Michael Seltzer. 2011 · 2011
Earlier work this paper cites.
The Uniform Effect of K-means Clustering , pages 17–35. Springer Berlin Heidelberg, Berlin, Heidelberg
Junjie Wu. 2012 · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Alternative structures for character-level rnns
Piotr Bojanowski, Armand Joulin, and Tomás Mikolov. 2015 · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Discriminative acoustic word embeddings: Recurrent neural network-based approaches
Shane Settle and Karen Livescu. 2016 · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
The zero resource speech challenge 2015: Proposed approaches and results
Maarten Versteegh, Xavier Anguera, Aren Jansen, and Emmanuel Dupoux. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Unsupervised Acoustic Unit Representation Learning for Voice Conversion Using WaveNet Auto-Encoders
Mingjie Chen and Thomas Hain. 2020 · 2020
Later among the works it cites.
A correspondence variational autoencoder for unsupervised acoustic word embeddings
Puyuan Peng, Herman Kamper, and Karen Livescu. 2020 · 2020
Later among the works it cites.
A comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings
Lisa Van Staden and Herman Kamper. 2020 · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The lj speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu. 2017 · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Yu-An Chung and James R. Glass. 2018 · 2018
Cited alongside, same era.
Herman Kamper. 2018 · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Cited alongside, same era.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Acoustic word embeddings for zero-resource languages using self-supervised contrastive learning and multilingual adaptation
Christiaan Jacobs, Yevgen Matusevych, and Herman Kamper. 2021 · 2021
Later among the works it cites.
Textless speech emotion conversion using decomposed and discrete representations
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu-Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. 2021 · 2021
Later among the works it cites.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Asr4real: An extended benchmark for speech models
Morgane Riviere, Jade Copet, and Gabriel Synnaeve. 2021 · 2021
Later among the works it cites.
Towards unsupervised learning of speech features in the wild
Morgane Rivière and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
SUPERB: Speech Processing Universal PERformance Benchmark
Shuwen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee. 2021 · 2021
Later among the works it cites.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021 · 2021
Later among the works it cites.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2022 · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022 · 2022
Later among the works it cites.
Fastdiff: A fast conditional diffusion model for high-quality speech synthesis
Rongjie Huang, Max W. Y. Lam, Jun Wang, Dan Su, Dong Yu, Yi Ren, and Zhou Zhao. 2022 · 2022
Later among the works it cites.
Word segmentation on discovered phone units with dynamic programming and self-supervised scoring
Herman Kamper. 2022 · 2022
Later among the works it cites.
textless-lib: a library for textless spoken language processing
Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Paden Tomasello, Ann Lee, Ali Elkahky, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, and Yossi Adi. 2022 · 2022
Later among the works it cites.
Reducing activation recomputation in large transformer models
Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2022 · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022 · 2022
Later among the works it cites.
Do coarser units benefit cluster prediction-based speech pre-training?
Ali Elkahky, Wei-Ning Hsu, Paden Tomasello, Tu-Anh Nguyen, Robin Algayres, Yossi Adi, Jade Copet, Emmanuel Dupoux, and Abdelrahman Mohamed. 2023 · 2023
Closest in time.
Word discovery in visually grounded, self-supervised speech models
Puyuan Peng and David Harwath. 2023 · 2023
Closest in time.