Fetching the paper…
Reading the bibliography…
Mixture-of-experts based acoustic models with dynamic routing mechanisms have proved promising results for speech recognition.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Network of experts for large-scale image categorization,”
Karim Ahmed, Mohammad Haris Baig, and Lorenzo Torresani, · 2016
Earlier work this paper cites.
“Hard mixtures of experts for large scale weakly supervised vision,”
Sam Gross, Marc’Aurelio Ranzato, and Arthur Szlam, · 2017
Earlier work this paper cites.
“Multi-scale dense networks for resource efficient image classification,”
Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Q Weinberger, · 2017
Earlier work this paper cites.
“Runtime neural pruning,”
Ji Lin, Yongming Rao, Jiwen Lu, and Jie Zhou, · 2017
Earlier work this paper cites.
“Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,”
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean, · 2017
Cited alongside, same era.
“Syllable-based acoustic modeling with ctc-smbr-lstm,”
Zhongdi Qu, Parisa Haghani, Eugene Weinstein, and Pedro Moreno, · 2017
Cited alongside, same era.
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
Bo Li, Tara N Sainath, Khe Chai Sim, Michiel Bacchiani, Eugene Weinstein, Patrick Nguyen, Zhifeng Chen, Yanghui Wu, and Kanishka Rao, · 2018
Cited alongside, same era.
“Gshard: Scaling giant models with conditional computation and automatic sharding,”
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen, · 2020
Cited alongside, same era.
“Deep mixture of experts via shallow embedding,”
Xin Wang, Fisher Yu, Lisa Dunlap, Yi-An Ma, Ruth Wang, Azalia Mirhoseini, Trevor Darrell, and Joseph E Gonzalez, · 2020
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
Tara N Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo-yiin Chang, Wei Li, Raziel Alvarez, Zhifeng Chen, et al., · 2020
Later among the works it cites.
“Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”
William Fedus, Barret Zoph, and Noam Shazeer, · 2021
Closest in time.
“Dynamic routing networks,”
Shaofeng Cai, Yao Shu, and Wei Wang, · 2021
Closest in time.
“Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts,”
Zhao You, Shulin Feng, Dan Su, and Dong Yu, · 2021
Closest in time.
“Fastmoe: A fast mixture-of-expert training system,”
Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, and Jie Tang, · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Closest in time.