Fetching the paper…
Reading the bibliography…
We present an open-source speech corpus for the Kazakh language.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
SRILM - An extensible language modeling toolkit
Andreas Stolcke. 2002 · 2002
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Kazakhstan-ethnicity, language and power
Bhavna Dave. 2007 · 2007
Earlier work this paper cites.
Cheap and fast - But is it good? Evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Y. Ng. 2008 · 2008
Earlier work this paper cites.
Cheap, fast and good enough: Automatic speech recognition with non-expert transcription
Scott Novotney and Chris Callison-Burch. 2010 · 2010
Earlier work this paper cites.
The Kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al. 2011 · 2011
Earlier work this paper cites.
Crowdsourcing for speech processing: Applications to data collection, transcription and assessment
Maxine Eskenazi, Gina-Anne Levow, Helen Meng, Gabriel Parent, and David Suendermann. 2013 · 2013
Earlier work this paper cites.
Assembling the kazakh language corpus
Olzhas Makhambetov, Aibek Makazhanov, Zhandos Yessenbayev, Bakhyt Matkarimov, Islam Sabyrgaliyev, and Anuar Sharafudinov. 2013 · 2013
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition
Awni Y. Hannun, C. Case, J. Casper, Bryan Catanzaro, Greg Diamos, E. Elsen, Ryan Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Ng. 2014 · 2014
Earlier work this paper cites.
Automatic Speech Recognition: A Deep Learning Approach
Dong Yu and Li Deng. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Cross-lingual transfer learning during supervised training in low resource scenarios
Amit Das and Mark Hasegawa-Johnson. 2015 · 2015
Cited alongside, same era.
On using monolingual corpora in neural machine translation
Çaglar Gülçehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loïc Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
A bilingual kazakh-russian system for automatic speech recognition and synthesis
Olga Khomitsevich, Valentin Mendelev, Natalia A. Tomashenko, Sergey Rybin, Ivan Medennikov, and Saule Kudubayeva. 2015 · 2015
Cited alongside, same era.
Audio augmentation for speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
AISHELL-2: transforming mandarin ASR research into industrial scale
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu. 2018 · 2018
Later among the works it cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Later among the works it cites.
Semi-orthogonal low-rank matrix factorization for deep neural networks
Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, Hainan Xu, Mahsa Yarmohammadi, and Sanjeev Khudanpur. 2018 · 2018
Later among the works it cites.
No need for a lexicon? Evaluating the value of the pronunciation lexica in end-to-end models
Tara N. Sainath, Rohit Prabhavalkar, Shankar Kumar, Seungji Lee, Anjuli Kannan, David Rybach, Vlad Schogol, Patrick Nguyen, Bo Li, Yonghui Wu, Zhifeng Chen, and Chung-Cheng Chiu. 2018 · 2018
Later among the works it cites.
CPJD corpus: Crowdsourced parallel speech corpus of Japanese dialects
Shinnosuke Takamichi and Hiroshi Saruwatari. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Purely sequence-trained neural networks for ASR based on lattice-free MMI
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Joint ctc-attention based end-to-end speech recognition using multi-task learning
Suyoun Kim, Takaaki Hori, and Shinji Watanabe. 2017 · 2017
Cited alongside, same era.
A free Kazakh speech database and a speech recognition baseline
Ying Shi, Askar Hamdullah, Zhiyuan Tang, Dong Wang, and Thomas Fang Zheng. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai. 2018 · 2018
Later among the works it cites.
Building the singapore english national speech corpus
Jia Xin Koh, Aqilah Mislan, Kevin Khoo, Brian Ang, Wilson Ang, Charmaine Ng, and Ying-Ying Tan. 2019 · 2019
Later among the works it cites.
Automatic recognition of kazakh speech using deep neural networks
Orken J. Mamyrbayev, Mussa Turdalyuly, Nurbapa Mekebayev, Keylan Alimhan, Aizat Kydyrbekova, and Tolganay Turdalykyzy. 2019 · 2019
Later among the works it cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
End-to-end speech recognition in agglutinative languages
Orken Mamyrbayev, Keylan Alimhan, Bagashar Zhumazhanov, Tolganay Turdalykyzy, and Farida Gusmanova. 2020 · 2020
Closest in time.
The rwth asr system for ted-lium release 2: Improving hybrid hmm with specaugment
Wei Zhou, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlüter, and Hermann Ney. 2020 · 2020
Closest in time.