Fetching the paper…
Reading the bibliography…
Self-supervised models for speech processing form representational spaces without using any external labels.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le. 2019 · 1904
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber. 2019 · 1912
Earlier work this paper cites.
Perception and production of syllable-initial english/r/and/l/by native speakers of japanese
Reiko A Yamada and Yoh’ichi Tohkura. 1990 · 1990
Earlier work this paper cites.
Speech perception in infancy predicts language development in the second year of life: A longitudinal study
Feng-Ming Tsao, Huei-Mei Liu, and Patricia K Kuhl. 2004 · 2004
Earlier work this paper cites.
Early speech perception and later language development: Implications for the" critical period"
Patricia K Kuhl, Barbara T Conboy, Denise Padden, Tobey Nelson, and Jessica Pruitt. 2005 · 2005
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2006
Earlier work this paper cites.
Infants show a facilitation effect for native language phonetic perception between 6 and 12 months
Patricia K Kuhl, Erica Stevens, Akiko Hayashi, Toshisada Deguchi, Shigeru Kiritani, and Paul Iverson. 2006 · 2006
Earlier work this paper cites.
The interspeech 2008 consonant challenge
MP Cooke and OE Scharenborg. 2008 · 2008
Earlier work this paper cites.
Evaluating computational models of infant phonetic learning across languages
Yevgen Matusevych, Thomas Schatz, Herman Kamper, Naomi H Feldman, and Sharon Goldwater. 2020 · 2008
Earlier work this paper cites.
On the assimilation-discrimination relationship in american english adults’ french vowel learning
Erika S Levy. 2009 · 2009
Earlier work this paper cites.
Human phoneme recognition depending on speech-intrinsic variability
Bernd T Meyer, Tim Jürgens, Thorsten Wesker, Thomas Brand, and Birger Kollmeier. 2010 · 2010
Earlier work this paper cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Yu Zhang, James Qin, Daniel S Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V Le, and Yonghui Wu. 2020 · 2010
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al. 2016 · 2016
Cited alongside, same era.
ABX-discriminability measures and applications
Thomas Schatz. 2016 · 2016
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. 2017 · 2017
Cited alongside, same era.
Unsupervised pretraining transfers well across languages
Morgane Riviere, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux. 2020 · 2020
Later among the works it cites.
Unsupervised pretraining transfers well across languages
Morgane Rivière, Armand Joulin, Pierre-Emmanuel Mazaré, and Emmanuel Dupoux. 2020 · 2020
Later among the works it cites.
The zero resource speech challenge 2021: Spoken language modelling
Ewan Dunbar, Mathieu Bernard, Nicolas Hamilakis, Tu Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Eugene Kharitonov, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Hubert: How much can a bad teacher benefit asr pre-training?
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021b · 2021
Later among the works it cites.
Generative spoken language modeling from raw audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Adelrahman Mohamed, et al. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visualizing phoneme category adaptation in deep neural networks
Odette Scharenborg, Sebastian Tiesmeyer, Mark Hasegawa-Johnson, and Najim Dehak. 2018 · 2018
Cited alongside, same era.
Neural network vs. hmm speech recognition systems as models of human cross-linguistic phonetic perception
Thomas Schatz and Naomi H Feldman. 2018 · 2018
Cited alongside, same era.
Metamers of neural networks reveal divergence from human perceptual systems
Jenelle Feather, Alex Durango, Ray Gonzalez, and Josh McDermott. 2019 · 2019
Cited alongside, same era.
Comparing unsupervised speech learning directly to human performance in speech perception
Juliette Millet, Nika Jurov, and Ewan Dunbar. 2019 · 2019
Cited alongside, same era.
Perceptimatic: A human speech perception benchmark for unsupervised subword modelling
Juliette Millet and Ewan Dunbar. 2020a · 2020
Cited alongside, same era.
The perceptimatic english benchmark for speech perception models
Juliette Millet and Ewan Dunbar. 2020b · 2020
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021a
Cited in the paper.
Later among the works it cites.
Predicting non-native speech perception using the perceptual assimilation model and state-of-the-art acoustic models
Juliette Millet, Ioana Chitoran, and Ewan Dunbar. 2021 · 2021
Later among the works it cites.
Towards unsupervised learning of speech features in the wild
Morgane Rivière and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Early phonetic learning without phonetic categories: Insights from large-scale simulations on realistic input
Thomas Schatz, Naomi H Feldman, Sharon Goldwater, Xuan-Nga Cao, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
The psychometrics of automatic speech recognition
Lotte Weerts, Stuart Rosen, Claudia Clopath, and Dan FM Goodman. 2021 · 2021
Later among the works it cites.
Self-training and pre-training are complementary for speech recognition
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli. 2021 · 2021
Later among the works it cites.