Fetching the paper…
Reading the bibliography…
The success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Darpa timit acoustic-phonetic continuous speech corpus cd-rom TIMIT, 1993
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, and N. Dahlgren · 1993
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol · 2008
Earlier work this paper cites.
An improved speech segmentation quality measure: the r-value
O. J. Räsänen, U. K. Laine, and T. Altosaar · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
A nonparametric bayesian approach to acoustic model discovery
C.-y. Lee and J. Glass · 2012
Earlier work this paper cites.
Japanese and korean voice search
M. Schuster and K. Nakajima · 2012
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
T. Mikolov, W.-t. Yih, and G. Zweig · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Variational inference for acoustic unit discovery
L. Ondel, L. Burget, and J. Černockỳ · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
R. Sennrich, B. Haddow, and A. Birch · 2016
Cited alongside, same era.
Hidden markov model variational autoencoder for acoustic unit discovery
J. Ebbers, J. Heymann, L. Drude, T. Glarner, R. Haeb-Umbach, and B. Raj · 2017
Cited alongside, same era.
A segmental framework for fully-unsupervised large-vocabulary speech recognition
H. Kamper, A. Jansen, and S. Goldwater · 2017
Cited alongside, same era.
Neural discrete representation learning
A. van den Oord, O. Vinyals, and k. kavukcuoglu · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Later among the works it cites.
Self-Supervised Contrastive Learning for Unsupervised Phoneme Segmentation
F. Kreuk, J. Keshet, and Y. Adi · 2020
Later among the works it cites.
Robust training of vector quantized bottleneck models
A. Łańcucki, J. Chorowski, G. Sanchez, R. Marxer, N. Chen, H. J. Dolfing, S. Khurana, T. Alumäe, and A. Laurent · 2020
Later among the works it cites.
The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling, 2020
T. A. Nguyen, M. de Seyssel, P. Rozé, M. Rivière, E. Kharitonov, A. Baevski, E. Dunbar, and E. Dupoux · 2020
Later among the works it cites.
Unsupervised pretraining transfers well across languages
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux · 2020
Later among the works it cites.
Unsupervised speech recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Unsupervised speech recognition via segmental empirical output distribution matching
C.-K. Yeh, J. Chen, C. Yu, and D. Yu · 2018
Cited alongside, same era.
Unsupervised neural segmentation and clustering for unit discovery in sequential data
J. Chorowski, N. Chen, R. Marxer, H. Dolfing, A. Łańcucki, G. Sanchez, T. Alumäe, and A. Laurent · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Putting an end to end-to-end: Gradient-isolated learning of representations
S. Löwe, P. O’Connor, and B. Veeling · 2019
Cited alongside, same era.
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli · 2021
Later among the works it cites.
Segmental Contrastive Predictive Coding for Unsupervised Word Segmentation
S. Bhati, J. Villalba, P. Żelasko, L. Moro-Velázquez, and N. Dehak · 2021
Later among the works it cites.
Aligned Contrastive Predictive Coding
J. Chorowski, G. Ciesielski, J. Dzikowski, A. Łańcucki, R. Marxer, M. Opala, P. Pusz, P. Rychlikowski, and M. Stypułkowski · 2021
Later among the works it cites.
Variable-rate discrete representation learning
S. Dieleman, C. Nash, J. Engel, and K. Simonyan · 2021
Later among the works it cites.
On generative spoken language modeling from raw audio
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, et al · 2021
Later among the works it cites.
Unsupervised speech segmentation and variable rate representation learning using segmental contrastive predictive coding
S. Bhati, J. Villalba, P. Żelasko, L. Moro-Velázquez, and N. Dehak · 2022
Closest in time.
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
S. Cuervo, M. Grabias, J. Chorowski, G. Ciesielski, A. Łańcucki, P. Rychlikowski, and R. Marxer · 2022
Closest in time.
Word segmentation on discovered phone units with dynamic programming and self-supervised scoring
H. Kamper · 2022
Closest in time.
Are discrete units necessary for Spoken Language Modeling?
T. A. Nguyen, B. Sagot, and E. Dupoux · 2022
Closest in time.