Fetching the paper…
Reading the bibliography…
Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios.
H. Hotelling, “Relations between two sets of variates,” Biometrika , vol. 28, no. 3/4, pp. 321–377, 1936
1936
Earlier work this paper cites.
T. Michaeli, W. Wang, and K. Livescu, “Nonparametric canonical correlation analysis,” in International Conference on Machine Learning . PMLR, 2016, pp. 1967–1976
1976
Earlier work this paper cites.
L. Rabiner and J. Wilpon, “Considerations in applying clustering techniques to speaker independent word recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 4, 1979, pp. 578–581
1979
Earlier work this paper cites.
D. K. Burton, J. E. Shore, and J. T. Buck, “A generalization of isolated word recognition using vector quantization,” Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 1983
1983
Earlier work this paper cites.
R. Gray, “Vector quantization,” IEEE ASSP Magazine , vol. 1, no. 2, pp. 4–29, 1984
1984
Earlier work this paper cites.
E. D. Petajan, “Automatic lipreading to enhance speech recognition (speech reading),” Ph.D. dissertation, University of Illinois at Urbana-Champaign, 1984
1984
Earlier work this paper cites.
J. Wilpon and L. Rabiner, “A modified k-means clustering algorithm for use in isolated work recognition,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 33, no. 3, pp. 587–594, 1985
1985
Earlier work this paper cites.
F. Soong, A. Rosenberg, and L. R. B. Juang, “A vector quantization approach to speaker recognition,” Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 1985
1985
Earlier work this paper cites.
L. Bahl, P. Brown, P. de Souza, and R. Mercer, “Maximum mutual information estimation of hidden Markov model parameters for speech recognition,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , vol. 11, 1986, pp. 49–52
1986
Earlier work this paper cites.
D. O’Shaughnessy, “Linear predictive coding,” IEEE Potentials , vol. 7, no. 1, pp. 29–32, 1988
1988
Earlier work this paper cites.
M. Legerstee, “Infants use multimodal information to imitate speech sounds,” Infant Behavior and Development , vol. 13, no. 3, pp. 343–354, 1990
1990
Earlier work this paper cites.
J. Westbury, P. Milenkovic, G. Weismer, and R. Kent, “X-ray microbeam speech production database,” JASA , vol. 88, no. S1, pp. S56–S56, 1990
1990
Earlier work this paper cites.
D. B. Paul and J. Baker, “The design for the Wall Street Journal-based CSR corpus,” in Speech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23–26 , 1992
1992
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 1992
1992
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning , vol. 8, no. 3, pp. 229–256, 1992
1992
Earlier work this paper cites.
M. Jordan and R. Jacobs, “Hierarchical mixtures of experts and the EM algorithm,” in Proceedings of 1993 International Conference on Neural Networks (IJCNN-93-Nagoya, Japan) , vol. 2, 1993
1993
Earlier work this paper cites.
J. S. Garofolo, “TIMIT Acoustic-Phonetic Continuous Speech Corpus LDC93S1,” Linguistic Data Consortium , 1993. [Online]. Available: https://catalog.ldc.upenn.edu/LDC93S1
1993
Earlier work this paper cites.
G. E. Hinton and R. Zemel, “Autoencoders, minimum description length and Helmholtz free energy,” in Advances in Neural Information Processing Systems , vol. 6, 1994
1994
Earlier work this paper cites.
J.-L. Gauvain and C.-H. Lee, “Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains,” IEEE Transactions on Speech and Audio Processing , vol. 2, no. 2, pp. 291–298, 1994
1994
Earlier work this paper cites.
S. Young and P. Woodland, “State clustering in hidden Markov model-based continuous speech recognition,” Computer Speech & Language , vol. 8, no. 4, pp. 369–383, 1994
1994
Earlier work this paper cites.
B. Olshausen and D. Field, “Emergence of simple-cell receptive field properties by learning a sparse code for natural images,” Nature , vol. 381, pp. 607–609, June 1996
1996
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine Learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
F. Jelinek, Statistical Methods for Speech Recognition . MIT press, 1997
1997
Earlier work this paper cites.
T. Kemp and A. Waibel, “Unsupervised training of a speech recognizer: Recent experiments,” Proceedings of European Conference on Speech Communication and Technology , 1999
1999
Earlier work this paper cites.
D. D. Lee and H. S. Seung, “Learning the parts of objects by nonnegative matrix factorization,” Nature , vol. 401, pp. 788–791, 1999
1999
Earlier work this paper cites.
D. Roy, “Learning from sights and sounds: A computational model,” PhD Thesis, MIT Media Laboratory , 1999
1999
Earlier work this paper cites.
——, “A neural implementation of canonical correlation analysis,” Neural Networks , vol. 12, no. 10, pp. 1391–1397, 1999
1999
Earlier work this paper cites.
P. L. Lai and C. Fyfe, “Kernel and nonlinear canonical correlation analysis,” Int. J. Neural Syst. , vol. 10, no. 5, pp. 365–377, 2000
2000
Earlier work this paper cites.
N. Smith and M. Gales, “Speech recognition using SVMs,” in NIPS , 2001
2001
Earlier work this paper cites.
A. Wrench, “A new resource for production modelling in speech technology,” Proc. Institute of Acoustics , vol. 23, no. 3, pp. 207–218, 2001
2001
Earlier work this paper cites.
L. Lamel, J.-L. Gauvain, and G. Adda, “Lightly supervised and unsupervised acoustic model training,” Computer Speech & Language , 2002
2002
Earlier work this paper cites.
G. E. Hinton, “Training products of experts by minimizing contrastive divergence,” Neural Comput. , vol. 14, no. 8, p. 1771–1800, aug 2002
2002
Earlier work this paper cites.
L. Wiskott and T. J. Sejnowski, “Slow feature analysis: Unsupervised learning of invariances,” Neural Computation , vol. 14, no. 4, pp. 715–770, 2002
2002
Earlier work this paper cites.
V. Hozjan, Z. Kacic, A. Moreno, A. Bonafonte, and A. Nogueiras, “Interface databases: Design and collection of a multilingual emotional speech database,” in Proceedings of International Conference on Language Resources and Evaluation , 2002
2002
Earlier work this paper cites.
M. Mohri, F. Pereira, and M. Riley, “Weighted finite-state transducers in speech recognition,” Computer Speech & Language , vol. 16, no. 1, pp. 69–88, 2002
2002
Earlier work this paper cites.
V. Venkataramani, S. Chakrabartty, and W. Byrne, “Support vector machines for segmental minimum Bayes risk decoding of continuous speech,” in ASRU , 2003
2003
Earlier work this paper cites.
V. Wan and S. Renals, “SVMSVM: support vector machine speaker verification methodology,” in ICASSP , 2003, pp. 221–224
2003
Earlier work this paper cites.
M. Schultz and T. Joachims, “Learning a distance metric from relative comparisons,” in Advances in Neural Information Processing Systems , S. Thrun, L. Saul, and B. Schölkopf, Eds., 2003
2003
Earlier work this paper cites.
G. Potamianos, C. Neti, G. Gravier, A. Garg, and A. W. Senior, “Recent advances in the automatic recognition of audiovisual speech,” Proceedings of the IEEE , vol. 91, no. 9, pp. 1306–1326, 2003
2003
Earlier work this paper cites.
B. Lee, M. Hasegawa-Johnson, C. Goudeseune, S. Kamdar, S. Borys, M. Liu, and T. Huang, “AVICAR: Audio-visual speech corpus in a car environment,” in Eighth International Conference on Spoken Language Processing , 2004
2004
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The Fisher corpus: A resource for the next generations of speech-to-text,” in Proceedings of International Conference on Language Resources and Evaluation , vol. 4, 2004, pp. 69–71
2004
Earlier work this paper cites.
F. R. Bach and M. I. Jordan, “A probabilistic interpretation of canonical correlation analysis,” Department of Statistics, University of California, Berkeley, Tech. Rep. 688, 2005
2005
Earlier work this paper cites.
J. Ma, S. Matsoukas, O. Kimball, and R. Schwartz, “Unsupervised training on large amounts of broadcast news data,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2006
2006
Earlier work this paper cites.
Y. LeCun, S. Chopra, R. Hadsell, F. J. Huang et al. , “A tutorial on energy-based learning,” in Predicting Structured Data . MIT Press, 2006
2006
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
H. Lee, A. Battle, R. Raina, and A. Y. Ng, “Efficient sparse coding algorithms,” in Proceedings of the 19th International Conference on Neural Information Processing Systems . MIT Press, 2006, p. 801–808
2006
Earlier work this paper cites.
P. S. Aleksic and A. K. Katsaggelos, “Audio-visual biometrics,” Proceedings of the IEEE , vol. 94, no. 11, pp. 2025–2044, 2006
2006
Earlier work this paper cites.
Y. Liu, P. Fung, Y. Yang, C. Cieri, S. Huang, and D. Graff, “HKUST/MTS: A very large scale Mandarin telephone speech corpus,” in Proceedings of International Symposium on Chinese Spoken Language Processing , 2006
2006
Earlier work this paper cites.
G. E. Hinton, “Learning multiple layers of representation,” Trends in Cognitive Sciences , vol. 11, pp. 428–434, 2007
2007
Earlier work this paper cites.
M. Ranzato, Y.-L. Boureau, S. Chopra, and Y. LeCun, “A unified energy-based framework for unsupervised learning,” in Proc. Eleventh International Conference on Artificial Intelligence and Statistics , 2007, pp. 371–379
2007
Earlier work this paper cites.
S. H. Weinberger and S. Kunath, “Towards a typology of English accents,” AACL Abstract Book , vol. 104, 2009
2009
Earlier work this paper cites.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” Journal of Machine Learning Research , 2010
2010
Earlier work this paper cites.
G. S. Sivaram, S. K. Nemala, M. Elhilali, T. D. Tran, and H. Hermansky, “Sparse coding for speech recognition,” in 2010 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2010, pp. 4346–4349
2010
Earlier work this paper cites.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” International Conference on Artificial Intelligence and Statistics (AISTATS) , 2010
2010
Earlier work this paper cites.
N. Dehak, P. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 19(4), pp. 788–798, 2011
2011
Earlier work this paper cites.
N. Dehak, P. .Torres-Carrasquillo, D. Reynolds, and R. Dehak, “Language recognition via i-vectors and dimensionality reduction,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2011
2011
Earlier work this paper cites.
H. Jégou, M. Douze, and C. Schmid, “Product quantization for nearest neighbor search,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 33, no. 1, pp. 117–128, 2011
2011
Earlier work this paper cites.
S. Narayanan, E. Bresch, P. K. Ghosh, L. Goldstein, A. Katsamanis, Y. Kim, A. Lammert, M. Proctor, V. Ramanarayanan, and Y. Zhu, “A multimodal real-time MRI articulatory corpus for speech research,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2011
2011
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in Proceedings of the 28th International Conference on Machine Learning , 2011
2011
Earlier work this paper cites.
M. A. Carlin, S. Thomas, A. Jansen, and H. Hermansky, “Rapid evaluation of speech representations for spoken term discovery,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2011
2011
Earlier work this paper cites.
S. Ravi and K. Knight, “Deciphering foreign language,” in Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies . Portland, Oregon, USA: Association for Computational Linguistics, Jun. 2011, pp. 12–21. [Online]. Available: https://aclanthology.org/P11-1002
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
H. A. Bourlard and N. Morgan, Connectionist speech recognition: A hybrid approach . Springer Science & Business Media, 2012, vol. 247
2012
Earlier work this paper cites.
M. U. Gutmann and A. Hyvärinen, “Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics.” Journal of Machine Learning Research , vol. 13, no. 2, 2012
2012
Earlier work this paper cites.
N. Srivastava and R. R. Salakhutdinov, “Multimodal learning with deep Boltzmann machines,” in Proceedings of Advances in Neural Information Processing Systems , 2012, pp. 2222–2230
2012
Earlier work this paper cites.
A. L. Maas, S. D. Miller, T. M. O’neil, A. Y. Ng, and P. Nguyen, “Word-level acoustic modeling with convolutional vector regression,” in Proc. ICML Workshop Representation Learn. , 2012
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Esteve, “TED-LIUM: An automatic speech recognition dedicated corpus,” in Proceedings of International Conference on Language Resources and Evaluation , 2012, pp. 125–129
2012
Earlier work this paper cites.
H. Gelas, L. Besacier, and F. Pellegrino, “Developments of Swahili resources for an automatic speech recognition system,” in Spoken Language Technologies for Under-Resourced Languages , 2012
2012
Earlier work this paper cites.
Y. Bengio, A. C. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Transactions on Pattern Analysis Machine Intelligence , vol. 35, no. 8, pp. 1798–1828, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. D. Zeiler, M. Ranzato, R. Monga, M. Mao, K. Yang, Q. V. Le, P. Nguyen, A. Senior, V. Vanhoucke, J. Dean et al. , “On rectified linear units for speech processing,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 3517–3521
2013
Earlier work this paper cites.
G. Andrew, R. Arora, J. A. Bilmes, and K. Livescu, “Deep canonical correlation analysis,” in Proceedings of the 30th International Conference on Machine Learning , 2013
2013
Earlier work this paper cites.
P.-S. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. Heck, “Learning deep structured semantic models for web search using clickthrough data,” in Int. Conf. Information and Knowledge Management , 2013
2013
Earlier work this paper cites.
K. Levin, K. Henry, A. Jansen, and K. Livescu, “Fixed-dimensional acoustic embeddings of variable-length segments in low-resource settings,” in Proceedings of IEEE Workshop on Automatic Speech Recognition and Understanding , 2013
2013
Earlier work this paper cites.
A. Jansen, E. Dupoux, S. Goldwater, M. Johnson, S. Khudanpur, K. Church, N. Feldman, H. Hermansky, F. Metze, R. Rose et al. , “A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2013, pp. 8111–8115
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14–16, 2014, Conference Track Proceedings , 2014
2014
Earlier work this paper cites.
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic backpropagation and approximate inference in deep generative models,” in International Conference on Machine Learning . PMLR, 2014, pp. 1278–1286
2014
Earlier work this paper cites.
L. Badino, C. Canevari, L. Fadiga, and G. Metta, “An auto-encoder based approach to unsupervised learning of subword units,” in 2014 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2014, pp. 7634–7638
2014
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov, “Dropout: A Simple Way to Prevent Neural Networks from Overfitting,” Journal of Machine Learning Research , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
G. Synnaeve, M. Versteegh, and E. Dupoux, “Learning words from images and speech,” in NIPS Workshop Learn. Semantics , 2014
2014
Earlier work this paper cites.
S. Bengio and G. Heigold, “Word embeddings for speech recognition,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2014
2014
Earlier work this paper cites.
M. J. Gales, K. M. Knill, A. Ragni, and S. P. Rath, “Speech recognition and keyword spotting for low-resource languages: BABEL project research at CUED,” in Fourth International Workshop on Spoken Language Technologies for Under-resourced Languages (SLTU-2014) , 2014
2014
Earlier work this paper cites.
M. Y. Tachbelie, S. T. Abate, and L. Besacier, “Using different acoustic, lexical and language modeling units for ASR of an under-resourced language–Amharic,” Speech Communication , vol. 56, pp. 181–194, 2014
2014
Earlier work this paper cites.
L. J. Rodríguez-Fuentes, A. Varona, M. Penagarikano, G. Bordel, and M. Diez, “GTTS-EHU systems for QUESST at MediaEval 2014,” in MediaEval , 2014
2014
Earlier work this paper cites.
M. Faruqi and C. Dyer, “Community evaluation and exchange of word vectors and wordvectors.org,” in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstration , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proceedings of Advances in Neural Information Processing Systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
M. I. Jordan and T. M. Mitchell, “Machine learning: Trends, perspectives, and prospects,” Science , vol. 349, no. 6245, pp. 255–260, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Cui, B. Kingsbury, B. Ramabhadran, A. Sethy, K. Audhkhasi, X. Cui, E. Kislal, L. Mangu, M. Nußbaum-Thom, M. Picheny, Z. Tüske, P. Golik, R. Schlüter, H. Ney, M. J. F. Gales, K. Knill, A. Ragni, H. Wang, and P. C. Woodland, “Multilingual representations for low resource speech recognition and keyword search,” 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015
2015
Earlier work this paper cites.
J. Chung, K. Kastner, L. Dinh, K. Goel, A. C. Courville, and Y. Bengio, “A Recurrent Latent Variable Model for Sequential Data,” in Proceedings of the 29th Conference on Neural Information Processing Systems (NeurIPS) , Montréal, Quebec, Canada, 2015, p. 9
2015
Earlier work this paper cites.
L. Badino, A. Mereta, and L. Rosasco, “Discovering discrete subword units with binarized autoencoders and hidden-markov-model encoders,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
H. Kamper, M. Elsner, A. Jansen, and S. Goldwater, “Unsupervised neural network based feature extraction using weak top-down constraints,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2015
2015
Earlier work this paper cites.
D. Renshaw, H. Kamper, A. Jansen, and S. Goldwater, “A comparison of neural network methods for unsupervised representation learning on the Zero Resource Speech Challenge,” Proceedings of the Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
R. B. Girshick, “Fast R-CNN,” in IEEE International Conference on Computer Vision , 2015
2015
Earlier work this paper cites.
W. Wang, R. Arora, K. Livescu, and J. A. Bilmes, “On deep multi-view representation learning,” in ICML , 2015
2015
Earlier work this paper cites.
W. Wang, R. Arora, K. Livescu, and J. Bilmes, “Unsupervised learning of acoustic features via deep canonical correlation analysis,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2015, pp. 4590–4594
2015
Earlier work this paper cites.
D. Harwath and J. R. Glass, “Deep multimodal semantic embeddings for speech and images,” in Proceedings of IEEE Workshop on Automatic Speech Recognition and Understanding , 2015
2015
Earlier work this paper cites.
K. Levin, A. Jansen, and B. Van Durme, “Segmental acoustic indexing for zero resource keyword search,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2015
2015
Earlier work this paper cites.
G. Chen, C. Parada, and T. N. Sainath, “Query-by-example keyword spotting using long short-term memory networks,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2015
2015
Earlier work this paper cites.
M. Ravanelli, L. Cristoforetti, R. Gretter, M. Pellin, A. Sosi, and M. Omologo, “The DIRHA-English corpus and related tasks for distant-speech recognition in domestic environments,” in Proceedings of IEEE Workshop on Automatic Speech Recognition and Understanding , 2015
2015
Earlier work this paper cites.
X. Anguera, L.-J. Rodriguez-Fuentes, A. Buzo, F. Metze, I. Szöke, and M. Penagarikano, “QUESST2014: Evaluating query-by-example speech search in a zero-resource setting with real-life queries,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Versteegh, R. Thiolliere, T. Schatz, X. N. Cao, X. Anguera, A. Jansen, and E. Dupoux, “The Zero Resource Speech Challenge 2015,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Zhang, P. Isola, and A. A. Efros, “Colorful image colorization,” CoRR , vol. abs/1603.08511, 2016
2016
Earlier work this paper cites.
M. Fraccaro, S. K. Sønderby, U. Paquet, and O. Winther, “Sequential Neural Models with Stochastic Layers,” in Proceedings of the 30th Conference on Neural Information Processing Systems (NeurIPS) , Barcelona, Spain, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y.-A. Chung, C.-C. Wu, C.-H. Shen, H.-Y. Lee, and L.-S. Lee, “Audio word2Vec: Unsupervised learning of audio segment representations using sequence-to-sequence autoencoder,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2016
2016
Earlier work this paper cites.
J. S. Chung and A. Zisserman, “Lip reading in the wild,” in Asian Conference on Computer Vision , 2016
2016
Earlier work this paper cites.
L. Badino, C. Canevari, L. Fadiga, and G. Metta, “Integrating articulatory data in deep neural network-based acoustic modeling,” Comp. Sp. & Lang. , vol. 36, pp. 173–195, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
D. Harwath, A. Torralba, and J. R. Glass, “Unsupervised learning of spoken language with visual context,” in Proceedings of Advances in Neural Information Processing Systems , 2016
2016
Earlier work this paper cites.
H. Kamper, W. Wang, and K. Livescu, “Deep convolutional acoustic word embeddings using word-pair side information,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2016
2016
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and K. MacDonald, “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,” 2016
2016
Earlier work this paper cites.
F. A. A. Laleye, L. Besacier, E. C. Ezin, and C. Motamed, “First automatic Fongbe continuous speech recognition system: Development of acoustic models and language models,” in FedCSIS . IEEE, 2016, pp. 477–482
2016
Earlier work this paper cites.
E. Gauthier, L. Besacier, S. Voisin, M. Melese, and U. P. Elingui, “Collecting resources in sub-Saharan African languages for automatic speech recognition: A case study of Wolof,” in LREC 2016 , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
R. Sennrich, B. Haddow, and A. Birch, “Improving neural machine translation models with monolingual data,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2016, pp. 86–96
2016
Cited alongside, same era.
L. Ondel, L. Burget, and J. Černockỳ, “Variational inference for acoustic unit discovery,” Procedia Computer Science , vol. 81, pp. 80–86, 2016
2016
Cited alongside, same era.
2020
Later among the works it cites.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-task self-supervised learning for robust speech recognition,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2020
2020
Later among the works it cites.
M. Tagliasacchi, B. Gfeller, F. de Chaumont Quitry, and D. Roblek, “Pre-training audio representations with self-supervision,” IEEE Signal Processing Letters , vol. 27, pp. 600–604, 2020
2020
Later among the works it cites.
M. Riviere, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2016
Cited alongside, same era.
W.-N. Hsu, Y. Zhang, and J. Glass, “Learning latent representations for speech generation and transformation,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2017
2017
Cited alongside, same era.
——, “Unsupervised learning of disentangled and interpretable representations from sequential data,” in Proceedings of Advances in Neural Information Processing Systems , 2017
2017
Cited alongside, same era.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” 2017
2017
Cited alongside, same era.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with Gumbel-softmax,” in Proceedings of International Conference on Learning Representations , 2017
2017
Cited alongside, same era.
G. Chrupała, L. Gelderloos, and A. Alishahi, “Representations of language in a model of visually grounded speech signal,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2017, pp. 613–622
2017
Cited alongside, same era.
S. Settle, K. Levin, H. Kamper, and K. Livescu, “Query-by-example search with discriminative neural acoustic word embeddings,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2017
2017
Cited alongside, same era.
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord, “Learning robust and multilingual speech representations,” in EMNLP , 2020
2020
Later among the works it cites.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in Proceedings of International Conference on Learning Representations , 2020
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Proceedings of Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
L. Wang and M. Hasegawa-Johnson, “A DNN-HMM-DNN hybrid model for discovering word-like units from spoken captions and image regions,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
P. Peng, H. Kamper, and K. Livescu, “A correspondence variational autoencoder for unsupervised acoustic word embeddings,” in Proc. NeurIPS Workshop on Self-Supervised Learning for Speech and Audio Processing , 2020
2020
Later among the works it cites.
S. Toshniwal, H. Shi, B. Shi, L. Gao, K. Livescu, and K. Gimpel, “A cross-task analysis of text span representations,” in Proceedings of the 5th Workshop on Representation Learning for NLP , 2020, pp. 166–176
2020
Later among the works it cites.
J. Kahn et al. , “Libri-light: A benchmark for ASR with limited or no supervision,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2020
2020
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A massively-multilingual speech corpus,” in Proceedings of International Conference on Language Resources and Evaluation , 2020
2020
Later among the works it cites.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A large-scale multilingual dataset for speech research,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
S. Khurana, A. Laurent, W.-N. Hsu, J. Chorowski, A. Lancucki, R. Marxer, and J. Glass, “A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
Y. Zhang, J. Qin, D. S. Park, W. Han, C.-C. Chiu, R. Pang, Q. V. Le, and Y. Wu, “Pushing the limits of semi-supervised learning for automatic speech recognition,” in Workshop on Self-Supervised Learning for Speech and Audio Processing, NeurIPS , 2020
2020
Later among the works it cites.
E. Dunbar, J. Karadayi, M. Bernard, X.-N. Cao, R. Algayres, L. Ondel, L. Besacier, S. Sakti, and E. Dupoux, “The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de Chaumont Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, “Towards Learning a Universal Non-Semantic Representation of Speech,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
A. Tjandra, S. Sakti, and S. Nakamura, “Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
B. van Niekerk, L. Nortje, and H. Kamper, “Vector-Quantized Neural Networks for Acoustic Unit Discovery in the ZeroSpeech 2020 Challenge,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
S. Ling, J. Salazar, Y. Liu, and K. Kirchhoff, “BERTphone: Phonetically-aware encoder representations for utterance-level speaker and language recognition,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop , 2020, pp. 9–16
2020
Later among the works it cites.
S. wen Yang, A. T. Liu, and H. yi Lee, “Understanding self-attention of self-supervised audio Transformers,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
G. Wang, A. Rosenberg, Z. Chen, Y. Zhang, B. Ramabhadran, Y. Wu, and P. Moreno, “Improving speech recognition using consistent predictions on synthesized speech,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2020, pp. 7029–7033
2020
Later among the works it cites.
A. Laptev, R. Korostik, A. Svischev, A. Andrusenko, I. Medennikov, and S. Rybin, “You do not need more data: Improving end-to-end speech recognition by text-to-speech data augmentation,” in 2020 13th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI) . IEEE, 2020, pp. 439–444
2020
Later among the works it cites.
Y. Huang, L. He, W. Wei, W. Gale, J. Li, and Y. Gong, “Using personalized speech synthesis and neural language generator for rapid speaker adaptation,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2020, pp. 7399–7403
2020
Later among the works it cites.
R. Masumura, N. Makishima, M. Ihori, A. Takashima, T. Tanaka, and S. Orihashi, “Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for End-to-End Automatic Speech Recognition,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill et al. , “On the opportunities and risks of foundation models,” 2021
2021
Later among the works it cites.
L. Ericsson, H. Gouk, C. C. Loy, and T. M. Hospedales, “Self-supervised representation learning: Introduction, advances and challenges,” 2021
2021
Later among the works it cites.
X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self-supervised learning: Generative or contrastive,” IEEE Transactions on Knowledge & Data Engineering , no. 01, pp. 1–1, Jun 2021
2021
Later among the works it cites.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” 2021
2021
Later among the works it cites.
L. Jing and Y. Tian, “Self-supervised visual feature learning with deep neural networks: A survey,” IEEE Transactions on Pattern Analysis & Machine Intelligence , vol. 43, no. 11, pp. 4037–4058, nov 2021
2021
Later among the works it cites.
S. Latif, R. Rana, S. Khalifa, R. Jurdak, J. Qadir, and B. W. Schuller, “Deep representation learning in speech processing: Challenges, recent advances, and future trends,” 2021
2021
Later among the works it cites.
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
L. Girin, S. Leglaive, X. Bie, J. Diard, T. Hueber, and X. Alameda-Pineda, “Dynamical variational autoencoders: A comprehensive review,” Foundations and Trends® in Machine Learning , vol. 15, pp. 1–175, 12 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Bell, J. Fainberg, O. Klejch, J. Li, S. Renals, and P. Swietojanski, “Adaptation algorithms for neural network-based speech recognition: An overview,” IEEE Open Journal of Signal Processing , vol. 2, pp. 33–66, 2021
2021
Later among the works it cites.
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-wise analysis of a self-supervised speech representation model,” in Proceedings of IEEE Workshop on Automatic Speech Recognition and Understanding , 2021
2021
Later among the works it cites.
A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W.-N. Hsu, A. Mohamed, and E. Dupoux, “Speech resynthesis from discrete disentangled self-supervised representations,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, and E. Dupoux, “On generative spoken language modeling from raw audio,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 1336–1354, 2021
2021
Later among the works it cites.
D. Jiang, W. Li, R. Zhang, M. Cao, N. Luo, Y. Han, W. Zou, K. Han, and X. Li, “A further study of unsupervised pretraining for Transformer based speech recognition,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2021
2021
Later among the works it cites.
X. Yue and H. Li, “Phonetically motivated self-supervised speech representation learning,” Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
A. T. Liu, S.-W. Li, and H.-y. Lee, “TERA: Self-supervised learning of Transformer encoder representation for speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2351–2366, 2021
2021
Later among the works it cites.
A. H. Liu, Y.-A. Chung, and J. Glass, “Non-Autoregressive Predictive Coding for Learning Speech Representations from Local Dependencies,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
J. Luo, J. Wang, N. Cheng, and J. Xiao, “Dropout regularization for self-supervised learning of Transformer encoder speech representation,” Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
P.-H. Chi, P.-H. Chung, T.-H. Wu, C.-C. Hsieh, S.-W. Li, and H.-y. Lee, “Audio ALBERT: A lite BERT for self-supervised learning of audio representation,” Proceedings of IEEE Spoken Language Technology Workshop , 2021
2021
Later among the works it cites.
S. Sadhu, D. He, C.-W. Huang, S. H. Mallidi, M. Wu, A. Rastrow, A. Stolcke, J. Droppo, and R. Maas, “wav2vec-C: A Self-Supervised Model for Speech Representation Learning,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
Y. Chung, Y. Zhang, W. Han, C. Chiu, J. Qin, R. Pang, and Y. Wu, “W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,” 2021
2021
Later among the works it cites.
D. Jiang, W. Li, M. Cao, W. Zou, and X. Li, “Speech SIMCLR: Combining contrastive and reconstruction objective for self-supervised speech representation learning,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, and F. Wei, “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” 2021
2021
Later among the works it cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision Transformers,” in IEEE International Conference on Computer Vision , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Rouditchenko, A. Boggust, D. Harwath, B. Chen, D. Joshi, S. Thomas, K. Audhkhasi, H. Kuehne, R. Panda, R. Feris et al. , “AVLnet: Learning audio-visual language representations from instructional videos,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
R. Sanabria, A. Waters, and J. Baldridge, “Talk, don’t write: A study of direct speech-based image retrieval,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
A. Bapna, Y. an Chung, N. Wu, A. Gulati, Y. Jia, J. H. Clark, M. Johnson, J. Riesa, A. Conneau, and Y. Zhang, “SLAM: A unified encoder for speech and language modeling via speech-text joint pre-training,” 2021
2021
Later among the works it cites.
S. Wang, L. Thompson, and M. Iyyer, “Phrase-BERT: Improved phrase embeddings from BERT with an application to corpus exploration,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 10 837–10 851
2021
Later among the works it cites.
L. van Staden and H. Kamper, “A comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings,” in Proceedings of IEEE Spoken Language Technology Workshop , 2021, pp. 927–934
2021
Later among the works it cites.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in ACL , 2021
2021
Later among the works it cites.
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised cross-lingual speech representation learning at scale,” 2021
2021
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, Z. Chen, Z. Chen, S. Liu, J. Wu, Y. Qian, F. Wei, J. Li, and X. Yu, “UniSpeech-SAT: Universal speech representation learning with speaker aware pre-training,” 2021
2021
Later among the works it cites.
G. Chen, S. Chai, G.-B. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Z. You, and Z. Yan, “GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
J. Valk and T. Alumäe, “VoxLingua107: A dataset for spoken language recognition,” in Proceedings of IEEE Spoken Language Technology Workshop , 2021
2021
Later among the works it cites.
T. Likhomanenko, Q. Xu, J. Kahn, G. Synnaeve, and R. Collobert, “slimIPL: Language-Model-Free Iterative Pseudo-Labeling,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
A. Hajavi and A. Etemad, “Siamese capsule network for end-to-end speaker recognition in the wild,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Evain, H. Nguyen, H. Le, M. Z. Boito, S. Mdhaffar, S. Alisamir, Z. Tong, N. Tomashenko, M. Dinarelli, T. Parcollet, A. Allauzen, Y. Estève, B. Lecouteux, F. Portet, S. Rossato, F. Ringeval, D. Schwab, and L. Besacier, “LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
B. van Niekerk, L. Nortje, M. Baas, and H. Kamper, “Analyzing speaker information in self-supervised models to improve zero-resource speech processing,” Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
C. Wang et al. , “Self-supervised learning for speech recognition with intermediate layer supervision,” arXiv e-print 2112.08778 , 2021
2021
Later among the works it cites.
Y.-A. Chung, Y. Belinkov, and J. Glass, “Similarity analysis of self-supervised speech representations,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2021
2021
Later among the works it cites.
J. Pu, Y. Yang, R. Li, O. Elibol, and J. Droppo, “Scaling effect of self-supervised models,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
W.-N. Hsu, A. Sriram, A. Baevski, T. Likhomanenko, Q. Xu, V. Pratap, J. Kahn, A. Lee, R. Collobert, G. Synnaeve, and M. Auli, “Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
G.-T. Lin, C.-J. Hsu, D.-R. Liu, H.-Y. Lee, and Y. Tsao, “Analyzing the robustness of unsupervised speech recognition,” 2021
2021
Later among the works it cites.
T. Maekaku, X. Chang, Y. Fujita, L.-W. Chen, S. Watanabe, and A. Rudnicky, “Speech Representation Learning Combining Conformer CPC with Deep Cluster for the ZeroSpeech Challenge 2021,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Lee, H. Gong, P. Duquenne, H. Schwenk, P. Chen, C. Wang, S. Popuri, J. Pino, J. Gu, and W. Hsu, “Textless speech-to-speech translation on real data,” CoRR , 2021
2021
Later among the works it cites.
E. B. Zaken, S. Ravfogel, and Y. Goldberg, “BitFit: Simple parameter-efficient fine-tuning for Transformer-based masked language-models,” 2021
2021
Later among the works it cites.
D. Guo, A. Rush, and Y. Kim, “Parameter-efficient transfer learning with diff pruning,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Online: Association for Computational Linguistics, Aug. 2021, pp. 4884–4896. [Online]. Available: https://aclanthology.org/2021.acl-long.378
2021
Later among the works it cites.
C.-I. J. Lai, Y. Zhang, A. H. Liu, S. Chang, Y.-L. Liao, Y.-S. Chuang, K. Qian, S. Khurana, D. Cox, and J. Glass, “PARP: Prune, adjust and re-prune for self-supervised speech recognition,” in Proceedings of Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
H.-J. Chang, S. wen Yang, and H. yi Lee, “DistilHuBERT: Speech representation learning by layer-wise distillation of hidden-unit BERT,” 2021
2021
Later among the works it cites.
A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, “A Survey of Quantization Methods for Efficient Neural Network Inference,” Jun. 2021
2021
Later among the works it cites.
S. Cao, Y. Kang, Y. Fu, X. Xu, S. Sun, Y. Zhang, and L. Ma, “Improving streaming Transformer based ASR under a framework of self-supervised learning,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2021
2021
Later among the works it cites.
H.-S. Choi, J. Lee, W. Kim, J. H. Lee, H. Heo, and K. Lee, “Neural analysis and synthesis: Reconstructing speech from self-supervised representations,” in NeurIPS , 2021
2021
Later among the works it cites.
H. Wu, B. Zheng, X. Li, X. Wu, H. yi Lee, and H. Meng, “Characterizing the adversarial vulnerability of speech self-supervised learning,” 2021
2021
Later among the works it cites.
L. Borgholt, J. D. Havtorn, J. Edin, L. Maaløe, and C. Igel, “A brief overview of unsupervised neural speech representation learning,” in The 2nd Workshop on Self-supervised Learning for Audio and Speech Processing (AAAI-SAS-2022) , 2022
2022
Closest in time.
A. Baevski, W. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “data2vec: A general framework for self-supervised learning in speech, vision and language,” 2022
2022
Closest in time.
2022
Closest in time.
B. Shi, W.-N. Hsu, K. Lakhotia, and A. Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” in International Conference on Learning Representations , 2022
2022
Closest in time.
B. Shi, W.-N. Hsu, and A. Mohamed, “Robust self-supervised audio-visual speech recognition,” in Interspeech , 2022
2022
Closest in time.
P. Peng and D. Harwath, “Fast-slow transformer for visually grounding speech,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2022
2022
Closest in time.
D. M. Chan, S. Ghosh, D. Chakrabarty, and B. Hoffmeister, “Multi-modal pre-training for automated speech recognition,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2022
2022
Closest in time.
P. Peng and D. Harwath, “Self-supervised representation learning for speech using visual grounding and masked language modeling,” in AAAI SAS workshop , 2022
2022
Closest in time.
2022
Closest in time.
J. Turian, J. Shier, H. R. Khan, B. Raj, B. W. Schuller, C. J. Steinmetz, C. Malloy, G. Tzanetakis, G. Velarde, K. McNally, M. Henry, N. Pinto, C. Noufi, C. Clough, D. Herremans, E. Fonseca, J. Engel, J. Salamon, P. Esling, P. Manocha, S. Watanabe, Z. Jin, and Y. Bisk, “Hear: Holistic evaluation of audio representations,” in Proceedings of the NeurIPS 2021 Competitions and Demonstrations Track , vol. 176, 2022, pp. 125–145
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
T. A. Nguyen, E. Kharitonov, J. Copet, Y. Adi, W.-N. Hsu, A. Elkahky, P. Tomasello, R. Algayres, B. Sagot, A. Mohamed, and E. Dupoux, “Generative spoken dialogue language modeling,” 2022
2022
Closest in time.
G.-T. Lin, Y.-S. Chuang, H.-L. Chung, S.-w. Yang, H.-J. Chen, S. Dong, S.-W. Li, A. Mohamed, H.-y. Lee, and L.-s. Lee, “DUAL: Discrete spoken unit adaptive learning for textless spoken question answering,” CoRR , 2022
2022
Closest in time.
B. Thomas, S. Kessler, and S. Karout, “Efficient adapter transfer of self-supervised speech models for automatic speech recognition,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing , 2022
2022
Closest in time.
K.-W. Chang, W.-C. Tseng, S.-W. Li, and H. yi Lee, “SpeechPrompt: An exploration of prompt tuning on generative spoken language model for speech processing tasks,” Proceedings of the Annual Conference of the International Speech Communication Association , 2022
2022
Closest in time.
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient Transformers: A Survey,” Mar. 2022
2022
Closest in time.
K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “ContentVec: An improved self-supervised speech representation by disentangling speakers,” in Proceedings of International Conference on Machine Learning , 2022
2022
Closest in time.
D. M. Chan and S. Ghosh, “Content-context factorized representations for automated speech recognition,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2022
2022
Closest in time.
K. P. Huang, Y.-K. Fu, Y. Zhang, and H.-y. Lee, “Improving distortion robustness of self-supervised speech processing tasks with domain adaptation,” in Proceedings of the Annual Conference of the International Speech Communication Association , 2022
2022
Closest in time.
H. Wang, Y. Qian, X. Wang, Y. Wang, C. Wang, S. Liu, T. Yoshioka, J. Li, and D. Wang, “Improving noise robustness of contrastive speech representation learning with speech reconstruction,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6062–6066
2022
Closest in time.
Q.-S. Zhu, J. Zhang, Z.-Q. Zhang, M.-H. Wu, X. Fang, and L.-R. Dai, “A noise-robust self-supervised pre-training model based speech representation learning for automatic speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 3174–3178
2022
Closest in time.