Fetching the paper…
Reading the bibliography…
Highly performing deep neural networks come at the cost of computational complexity that limits their practicality for deployment on portable devices.
“Hkust/mts: A very large scale mandarin telephone speech corpus,”
Yi Liu, Pascale Fung, Yongsheng Yang, Christopher Cieri, Shudong Huang, and David Graff, · 2006
Earlier work this paper cites.
“Low-rank matrix factorization for deep neural network training with high-dimensional output targets,”
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“An empirical exploration of ctc acoustic models,”
Yajie Miao, Mohammad Gowayyed, Xingyu Na, Tom Ko, Florian Metze, and Alexander Waibel, · 2016
Earlier work this paper cites.
“Joint ctc-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Factorization tricks for lstm networks,”
Oleksii Kuchaiev and Boris Ginsburg, · 2017
Cited alongside, same era.
“Advances in joint ctc-attention based end-to-end speech recognition with a deep cnn encoder and rnn-lm,”
Takaaki Hori, Shinji Watanabe, Yu Zhang, and William Chan, · 2017
Cited alongside, same era.
“Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Cited alongside, same era.
“Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition,”
Linhao Dong, Shuang Xu, and Bo Xu, · 2018
Cited alongside, same era.
“Building competitive direct acoustics-to-word models for english conversational speech recognition,”
“The speechtransformer for large-scale mandarin chinese speech recognition,”
Jie Li, Xiaorui Wang, Yan Li, et al., · 2019
Closest in time.
“Shrinkml: End-to-end asr model compression using reinforcement learning,”
Łukasz Dudziak, Mohamed Abdelfattah, Ravichander Vipperla, Stefanos Laskaridis, and Nicholas Lane, · 2019
Closest in time.
“On the effectiveness of low-rank matrix factorization for lstm model compression,”
Genta Indra Winata, Andrea Madotto, Jamin Shin, Elham J Barezi, and Pascale Fung, · 2019
Closest in time.
“End-to-end speech recognition with adaptive computation steps,”
Mohan Li, Min Liu, and Hattori Masanori, · 2019
Closest in time.
“Framewise supervised training towards end-to-end speech recognition models: First results,”
Mohan Li, Yuanjiang Cao, Weicong Zhou, and Min Liu, · 2019
Closest in time.
“Code-switched language models using neural based synthetic data from parallel sentences,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, and Michael Picheny, · 2018
Cited alongside, same era.
Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, and Pascale Fung, · 2019
Closest in time.