Fetching the paper…
Reading the bibliography…
This paper proposes multistream CNN, a novel neural network architecture for robust acoustic modeling in speech recognition tasks.
“How do humans process and recognize speech?,”
John Allen, · 1994
Earlier work this paper cites.
“Connectionist speech recognition: A hybrid approach,”
Herve Bourlard and Nelson Morgan, · 1994
Earlier work this paper cites.
“Towards subband-based speech recognition,”
Herve Bourlard, Stephane Dupont, Hynek Hermansky, and Nathaniel Morgan, · 1996
Earlier work this paper cites.
“A new ASR approach based on independent processing and recombination of partial frequency bands,”
Herve Bourlard and Stephane Dupont, · 1996
Earlier work this paper cites.
“Towards ASR on partially corrupted speech,”
Hynek Hermansky, Sangita Tibrewala, and Misha Pavel, · 1996
Earlier work this paper cites.
“Subband-based speech recognition,”
Herve Bourlard and Stephane Dupont, · 1997
Earlier work this paper cites.
“Sub-band based recognition of noisy speech,”
Sangita Tibrewala and Hynek Hermansky, · 1997
Earlier work this paper cites.
“Multi-band speech recognition in noisy environments,”
Shigeki Okawa, Enrico Bocchieri, and Alexandros Potamianos, · 1998
Earlier work this paper cites.
“Temporal patterns (TRAPS) in ASR noisy speech,”
Hynek Hermansky and Sanjita Sharma, · 1999
Earlier work this paper cites.
“Multi-stream adaptive evidence combination for noise robust ASR,”
Andrew Morris, Astrid Hagen, Hervé Glotin, and Hervé Bourlard, · 2001
Earlier work this paper cites.
“Multi-resolution RASTA filtering for TANDEM-based ASR,”
Hynek Hermansky and Sanjita Sharma, · 2005
Cited alongside, same era.
“Adaptive stream fusion in multistream recognition of speech,”
Nima Mesgarani, Samuel Thomas, and Hynek Hermansky, · 2011
Cited alongside, same era.
“The Kaldi speech recognition toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely, · 2011
Cited alongside, same era.
“Autoencoder based multi-stream combination for noise robust speech recognition,”
Sri Harish Mallidi, Tetsuji Ogawa, Karel Vesely, Phani S. Nidadavolu, and Hynek Hermansky, · 2015
Cited alongside, same era.
“LibrSspeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Cited alongside, same era.
“Novel neural network based fusion for multistream ASR,”
“The third ‘CHIME’ speech separation and recognition challenge: Analysis and outcome,”
Jon Barker, Ricard Marxer, Emmanuel Vincent, and Shinji Watanabe, · 2017
Later among the works it cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Acoustic modeling of speech waveform based on multi-resolution neural network signal processing,”
Zoltan Tuske, Ralf Schluter, and Hermann Ney, · 2018
Later among the works it cites.
“Multi-encoder multi-resolution framework for end-to-end speech recognition,”
Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi, Takaaki Hori, Shinji Watanabe, and Hynek Hermansky, · 2018
Later among the works it cites.
“Semi-orthogonal low-rank matrix factorization for deep neural networks,”
Daniel Povey, Gaofeng Cheng, Yiming Wang, Ke Li, Hainan Xu, Mahsa Yarmohamadi, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sri Harish Mallidi and Hynek Hermansky, · 2016
Cited alongside, same era.
“A framework for practical multistream ASR,”
Sri Harish Mallidi and Hynek Hermansky, · 2016
Cited alongside, same era.
“Hierarchical attention networks for document classification,”
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy, · 2016
Cited alongside, same era.
“Realistic multi-microphone data simulation for distant speech recognition,”
Maurizio Omologo Mirco Ravanelli, Piergiorgio Svaizer, · 2016
Cited alongside, same era.
“Purely sequence-trained neural networks for ASR based on lattice-free MMI,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahrmani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Cited alongside, same era.
“Recurrent neural network language model adaptation for conversational speech recognition,”
Ke Li, Hainan Xu, Yiming Wang, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“Multi-stride self-attention for speech recognition,”
Kyu J. Han, Jing Huang, Yun Tang, Xiaodong He, and Bowen Zhou, · 2019
Later among the works it cites.
“State-of-the-art speech recognition using multi-stream self-attention with dilated 1D convolution,”
Kyu J. Han, Ramon Prieto, and Tao Ma, · 2019
Later among the works it cites.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiua, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Later among the works it cites.
“ASAPP-ASR: Multistream CNN and self-attentive SRU for SOTA speech recognition,”
Jing Pan, Joshua Shapiro, Jeremy Wohlwend, Kyu J. Han, Tao Lei, and Tao Ma, · 2020
Closest in time.