Fetching the paper…
Reading the bibliography…
In the last decade of automatic speech recognition (ASR) research, the introduction of deep learning brought considerable reductions in word error rate of more than 50% relative, compared to modeling without deep learning.
1904
Earlier work this paper cites.
1907
Earlier work this paper cites.
1907
Earlier work this paper cites.
1911
Earlier work this paper cites.
1911
Earlier work this paper cites.
1912
Earlier work this paper cites.
1912
Earlier work this paper cites.
P. Haffner, “Connectionist Speech Recognition with a Global MMI Algorithm,” in Proc. Eurospeech , Berlin, Germany, Dec. 1993, pp. 1929–1932
1932
Earlier work this paper cites.
G. K. Zipf, Human Behavior and the Principle of Least Effort . Boston, MA: Addison-Wesley Press, 1949
1949
Earlier work this paper cites.
R. E. Bellman, Dynamic Programming . Princeton, NJ: Princeton University Press, 1957
1957
Earlier work this paper cites.
B. Polyak, “Some Methods of Speeding up the Convergence of Iteration Methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964. [Online]. Available: https://www.sciencedirect.com/science/article/pii/0041555364901375
1964
Earlier work this paper cites.
A. Viterbi, “Error Bounds for Convolutional Codes and an Asymptotically Optimal Decoding Algorithm,” IEEE Transactions on Information Theory , vol. 13, pp. 260–269, 1967
1967
Earlier work this paper cites.
L. Baum, “An Inequality and Associated Maximization Technique in Statistical Estimation for Probabilistic Functions of Markov Processes,” Inequalities , vol. 3, pp. 1–8, 1972
1972
Earlier work this paper cites.
Y. Nesterov, “A method of solving a convex programming problem with convergence rate O( 1 k 2 \frac{1}{k^{2}} ),” Soviet Mathematics Doklady , vol. 27, pp. 372–376, 1983
1983
Earlier work this paper cites.
L. R. Bahl, F. Jelinek, and R. L. Mercer, “A Maximum Likelihood Approach to Continuous Speech Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 5, no. 2, pp. 179–190, Mar. 1983
1983
Earlier work this paper cites.
L. Breiman, J. Friedman, C. Stone, and R. Olshen, Classication and Regression Trees . Belmont, CA: Taylor & Francis, 1984
1984
Earlier work this paper cites.
H. Ney, “The Use of a One-Stage Dynamic Programming Algorithm for Connected Word Recognition,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 32, no. 2, pp. 263–271, 1984
1984
Earlier work this paper cites.
L. Rabiner and B.-H. Juang, “An Introduction to Hidden Markov Models,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 3, no. 1, pp. 4–16, 1986
1986
Earlier work this paper cites.
T. P. Vogl, J. Mangis, A. Rigler, W. Zink, and D. Alkon, “Accelerating the Convergence of the Back-Propagation Method,” Biological Cybernetics , vol. 59, no. 4, pp. 257–263, 1988
1988
Earlier work this paper cites.
M. Nakamura and K. Shikano, “A Study of English Word Category Prediction Based on Neural Networks,” in Proc. IEEE ICASSP , Glasglow, UK, May 1989, pp. 731–734
1989
Earlier work this paper cites.
R. J. Williams and J. Peng, “An Efficient Gradient-Based Algorithm for On-Line Training of Recurrent Network Trajectories,” IEEE Neural Computation , vol. 2, no. 4, pp. 490–501, 1990
1990
Earlier work this paper cites.
Y. Bengio, R. De Mori, G. Flammia, and R. Kompe, “Neural Network-Gaussian Mixture Hybrid for Speech Recognition or Density Estimation,” in Proc. NIPS , vol. 4, Colorado, Dec. 1991, pp. 175–182
1991
Earlier work this paper cites.
S. Renals, N. Morgan, H. Bourlard, C. Wooters, and P. Kohn, “Connectionist Speech Recognition: Status and Prospects,” ICSI, 1991, Tech. Rep. TR-OI-070
1991
Earlier work this paper cites.
A. Krogh and J. Hertz, “A Simple Weight Decay Can Improve Generalization,” in Neural Information Processing Systems (NIPS) , Denver, CO, Dec. 1991, pp. 950–957
1991
Earlier work this paper cites.
J. Godfrey, E. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone Speech Corpus for Research and Development,” in Proc. IEEE ICASSP , vol. 1, San Francisco, CA, Mar. 1992, pp. 517–520 vol.1
1992
Earlier work this paper cites.
H. A. Bourlard and N. Morgan, Connectionist Speech Recognition: a Hybrid Approach . Norwell, MA: Kluwer Academic Publishers, 1993
1993
Earlier work this paper cites.
J. L. Elman, “Learning and Development in Neural Networks: The Importance of Starting Small,” Cognition , vol. 48, no. 1, pp. 71–99, 1993. [Online]. Available: https://www.sciencedirect.com/science/article/pii/0010027793900584
1993
Earlier work this paper cites.
S. J. Young and P. C. Woodland, “The Use of State Tying in Continuous Speech Recognition,” in Proc. Eurospeech , Berlin, Germany, Dec. 1993, pp. 2203–2206
1993
Earlier work this paper cites.
A. F. Murray and P. J. Edwards, “Enhanced MLP Performance and Fault Tolerance Resulting from Synaptic Weight Noise during Training,” IEEE Transactions on Neural Networks , vol. 5, no. 5, pp. 792–802, Sep. 1994
1994
Earlier work this paper cites.
R. Haeb-Umbach and H. Ney, “Improvements in Beam Search for 10000-Word Continuous-Speech Recognition,” IEEE Transactions on Speech and Audio Processing , vol. 2, no. 2, pp. 353–356, 1994
1994
Earlier work this paper cites.
C. Wooters and A. Stolcke, “Multiple-Pronunciation Lexical Modeling in a Speaker Independent Speech Understanding System,” in Proc. ICSLP , Yokohama, Japan, Sep. 1994, pp. 1363–1366
1994
Earlier work this paper cites.
J. Makhoul and R. Schwartz, “State of the Art in Continuous Speech Recognition,” Proc. NAS , vol. 92, no. 22, pp. 9956–9963, Oct. 1995
1995
Earlier work this paper cites.
S. F. Chen and J. Goodman, “An Empirical Study of Smoothing Techniques for Language Modeling,” in Proc. ACL , Santa Cruz, CA, Jun. 1996, pp. 310–318
1996
Earlier work this paper cites.
F. Jelinek, Statistical Methods for Speech Recognition . Cambridge, MA: MIT Press, 1997
1997
Earlier work this paper cites.
V. Fontaine, C. Ris, and H. Leich, “Nonlinear Discriminant Analysis for Improved Speech Recognition,” in Proc. Eurospeech , Rhodes, Greece, Sep. 1997, pp. 1–4
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
V. Valtchev, J. J. Odell, P. C. Woodland, and S. J. Young, “MMIE Training of Large Vocabulary Recognition Systems,” Speech Communication , vol. 22, no. 4, pp. 303–314, 1997
1997
Earlier work this paper cites.
N. Deshmukh, A. Ganapathiraju, and J. Picone, “Hierarchical Search for Large-Vocabulary Conversational Speech Recognition: Working Toward a Solution to the Decoding Problem,” IEEE Signal Processing Magazine , vol. 16, no. 5, pp. 84–107, 1999
1999
Earlier work this paper cites.
L. Nguyen and R. Schwartz, “Single-Tree Method for Grammar-Directed Search,” in Proc. IEEE ICASSP , vol. 2, Phoenix, AZ, Mar. 1999, pp. 613–616
1999
Earlier work this paper cites.
H. Hermansky, D. Ellis, and S. Sharma, “Tandem connectionist Feature Extraction for Conventional HMM Systems,” in Proc. IEEE ICASSP , vol. 3, Istanbul, Turkey, Jun. 2000, pp. 1635–1638
2000
Earlier work this paper cites.
Y. Bengio, R. Ducharme, and P. Vincent, “A Neural Probabilistic Language Model,” in Proc. NIPS , vol. 13, Denver, CO, Nov. 2000, pp. 932–938
2000
Earlier work this paper cites.
H. Ney and S. Ortmanns, “Progress in Dynamic Programming Search for LVCSR,” Proceedings of the IEEE , vol. 88, no. 8, pp. 1224–1240, Aug. 2000. [Online]. Available: http://dx.doi.org/10.1109/5.880081
2000
Earlier work this paper cites.
D. Povey and P. Woodland, “Improved Discriminative Training Techniques for Large Vocabulary Continuous Speech Recognition,” in Proc. IEEE ICASSP , Salt Lake City, UT, May 2001, pp. 45–48
2001
Earlier work this paper cites.
R. Schlüter, W. Macherey, B. Müller, and H. Ney, “Comparison of Discriminative Training Criteria and Optimization Methods for Speech Recognition,” Speech Communication , vol. 34, no. 3, pp. 287–310, May 2001, EURASIP Best Paper Award
2001
Earlier work this paper cites.
H. Schwenk and J.-L. Gauvain, “Connectionist Language Modeling for Large Vocabulary Continuous Speech Recognition,” in Proc. IEEE ICASSP , Orlando, FL, May 2002, pp. 765–768
2002
Earlier work this paper cites.
M. Mohri, F. Pereira, and M. Riley, “Weighted Finite-State Transducers in Speech Recognition,” Computer Speech & Language , vol. 16, no. 1, pp. 69–88, 2002
2002
Earlier work this paper cites.
S. Kanthak and H. Ney, “Context-Dependent Acoustic Modeling Using Graphemes for Large Vocabulary Speech Recognition,” in Proc. IEEE ICASSP , Orlando, FL, May 2002, pp. 845–848
2002
Earlier work this paper cites.
D. Klakow and J. Peters, “Testing the Correlation of Word Error Rate and Perplexity,” Speech Communication , vol. 38, no. 1, pp. 19–28, 2002
2002
Earlier work this paper cites.
D. Johnson, D. Ellis, C. Oei, C. Wooters, and P. Faerber, “QuickNet,” ICSI, Berkeley, 2004. [Online]. Available: http://www.icsi.berkeley.edu/Speech/qn.html
2004
Earlier work this paper cites.
2005
Earlier work this paper cites.
H. Soltau, B. Kingsbury, L. Mangu, D. Povey, G. Saon, and G. Zweig, “The IBM 2004 Conversational Telephony System for Rich Transcription,” in Proc. IEEE ICASSP , Philadelphia, PA, Mar. 2005, pp. 205–208
2005
Earlier work this paper cites.
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , Pittsburgh, PA, Jun. 2006, pp. 369–376
2006
Earlier work this paper cites.
P. Liang, A. Bouchard-Côté, D. Klein, and B. Taskar, “An End-to-End Discriminative Approach to Machine Translation,” in Proc. ACL , Sydney, Australia, Jul. 2006, p. 761–768
2006
Earlier work this paper cites.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A Fast Learning Algorithm for Deep Belief Nets,” Neural Computation , vol. 18, no. 7, pp. 1527–1554, Jul. 2006
2006
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, “Greedy Layer-Wise Training of Deep Networks,” in Proc. NIPS , Barcelona, Spain, Dec. 2006, pp. 153–160
2006
Earlier work this paper cites.
2007
Earlier work this paper cites.
X. He, L. Deng, and W. Chou, “Discriminative Learning in Sequential Pattern Recognition – A Unifying Review for Optimization-Oriented Speech Recognition,” IEEE Signal Processing Magazine , vol. 25, no. 5, pp. 14–36, 2008
2008
Earlier work this paper cites.
S. Scanzio, P. Laface, L. Fissore, R. Gemello, and F. Mana, “On the Use of a Multilingual Neural Network Front-End,” in Proc. Interspeech , Brisbane, Australia, Sep. 2008, pp. 2711–2714
2008
Earlier work this paper cites.
B. Kingsbury, “Lattice-Based Optimization of Sequence Classification Criteria for Neural-Network Acoustic Modeling,” in Proc. IEEE ICASSP , Taipei, Taiwan, Apr. 2009, pp. 3761–3764
2009
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum Learning,” in Proc. ICML , Montreal, Quebec, Canada, Jun. 2009, p. 41–48
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur, “Recurrent Neural Network Based Language Model,” in Proc. Interspeech , Makuhari, Japan, Sep. 2010, pp. 1045–1048
2010
Earlier work this paper cites.
S. Wiesler, G. Heigold, M. Nußbaum-Thom, R. Schlüter, and H. Ney, “A Discriminative Splitting Criterion for Phonetic Decision Trees,” in Proc. Interspeech , Makuhari, Japan, Sep. 2010, pp. 54–57, one of shortlist for Best Student Paper Award
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
F. Seide, G. Li, and D. Yu, “Conversational Speech Transcription Using Context-Dependent Deep Neural Networks,” in Proc. Interspeech , Florence, Italy, Aug. 2011, pp. 437–440
2011
Earlier work this paper cites.
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa, “Natural Language Processing (Almost) from Scratch,” Journal of Machine Learning Research , vol. 12, pp. 2493–2537, 2011
2011
Earlier work this paper cites.
A. Graves, “Practical Variational Inference for Neural Networks,” Advances in Neural Information Processing Systems , vol. 24, 2011
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and Korean Voice Search,” in Proc. IEEE ICASSP , Kyoto, Japan, Mar. 2012, pp. 5149–5152
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
G. Heigold, R. Schlüter, H. Ney, and S. Wiesler, “Discriminative Training for Automatic Speech Recognition: Modeling, Criteria, Optimization, Implementation, and Performance,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 58–69, Nov. 2012
2012
Earlier work this paper cites.
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Advances in Neural Information Processing Systems (NIPS) , vol. 25, Lake Tahoe, NV, Dec. 2012
2012
Earlier work this paper cites.
A. Graves, “Connectionist Temporal Classification,” in Supervised Sequence Labelling with Recurrent Neural Networks . Heidelberg, Germany: Springer, 2012, ch. Connectionist Temporal Classification, pp. 61–93
2012
Earlier work this paper cites.
M. Sundermeyer, R. Schlüter, and H. Ney, “LSTM Neural Networks for Language Modeling,” in Proc. Interspeech , Portland, OR, Sep. 2012, pp. 194–197
2012
Earlier work this paper cites.
I. McGraw, I. Badr, and J. R. Glass, “Learning Lexicons From Speech Using a Pronunciation Mixture Model,” IEEE/ACM Trans. Audio, Speech, and Language Processing , vol. 21, no. 2, pp. 357–366, 2012
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech Recognition with Deep Recurrent Neural Networks,” in Proc. IEEE ICASSP , Vancouver, BC, Canada, May 2013, pp. 6645–6649
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Z. Tüske, J. Pinto, D. Willett, and R. Schlüter, “Investigation on Cross- and Multilingual MLP features under matched and mismatched acoustical conditions,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , Vancouver, Canada, May 2013, pp. 7349–7353
2013
Earlier work this paper cites.
A. Senior, G. Heigold, M. Ranzato, and K. Yang, “An Empirical Study of Learning Rates in Deep Neural Networks for Speech Recognition,” in Proc. IEEE ICASSP . Vancouver, BC, Canada: IEEE, May 2013, pp. 6724–6728
2013
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the Importance of Initialization and Momentum in Deep Learning,” in Proc. ICML , Atlanta, GA, Jun. 2013, pp. 1139–1147
2013
Earlier work this paper cites.
L. Wan, M. Zeiler, S. Zhang, Y. Le Cun, and R. Fergus, “Regularization of Neural Networks using DropConnect,” in Proc. ICML , 2013, pp. 1058–1066
2013
Earlier work this paper cites.
N. Kanda, R. Takeda, and Y. Obuchi, “Elastic Spectral Distortion for Low Resource Speech Recognition with Deep Neural Networks,” in Proc. IEEE ASRU , Olomouc, Czech Republic, Dec. 2013, pp. 309–314
2013
Earlier work this paper cites.
N. Jaitly and G. E. Hinton, “Vocal Tract Length Perturbation (VTLP) Improves Speech Recognition,” in Proc. ICML , vol. 117, Jun. 2013, p. 21
2013
Earlier work this paper cites.
T. Hori and A. Nakamura, Speech Recognition Algorithms Using Weighted Finite-State Transducers . San Rafael, CA: Morgan & Claypool Publishers, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Z. Tüske, P. Golik, R. Schlüter, and H. Ney, “Acoustic Modeling with Deep Neural Networks Using Raw Time Signal for LVCSR,” in Proc. Interspeech , Singapore, Sep. 2014, pp. 890–894
2014
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards End-to-End Speech Recognition with Recurrent Neural Networks,” in Proc. ICML , Beijing, China, Jun. 2014, pp. 1764–1772
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T. Hori, Y. Kubo, and A. Nakamura, “Real-Time One-Pass Decoding with Recurrent Neural Network Language Model for Speech Recognition,” in Proc. IEEE ICASSP , Florence, Italy, May 2014, pp. 6364–6368
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Z. Huang, G. Zweig, and B. Dumoulin, “Cache Based Recurrent Neural Network Language Model Inference for First Pass Speech Recognition,” in Proc. IEEE ICASSP , Florence, Italy, May 2014, pp. 6354–6358
2014
Earlier work this paper cites.
A. Senior, G. Heigold, M. Bacchiani, and H. Liao, “GMM-Free DNN Acoustic Model Training,” in Proc. IEEE ICASSP , Florence, Italy, May 2014, pp. 5602–5606. [Online]. Available: https://doi.org/10.1109/ICASSP.2014.6854675
2014
Earlier work this paper cites.
T. N. Sainath, R. J. Weiss, K. W. Wilson, A. Narayanan, M. Bacchiani, and A. Senior, “Speaker Location and Microphone Spacing Invariant Acoustic Modeling from Raw Multichannel Waveforms,” in Proc. IEEE ASRU , Scottsdale, AZ, Dec. 2015, pp. 30–36
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-Based Models for Speech Recognition,” in Proc. NIPS , vol. 28, Laval, Queèbec, Canada, Dec. 2015, pp. 577–585
2015
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural Machine Translation of Rare Words with Subword Units,” in Proc. ACL , Berlin, Germany, Aug. 2015, pp. 1715–1725
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
N. S. Keskar and G. Saon, “A Nonmonotone Learning Rate Strategy for SGD Training of Deep Neural Networks,” in Proc. IEEE ICASSP . Queensland, Australia: IEEE, Apr. 2015, pp. 4974–4978
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, “Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks,” Proc. NIPS , vol. 28, Dec. 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” in Proc. ICML , Lille, France, Jul. 2015, pp. 448–456
2015
Earlier work this paper cites.
T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio Augmentation for Speech Recognition,” in Proc. Interspeech , Dresden, Germany, Sep. 2015
2015
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-End Speech Recognition Using Deep RNN Models and WFST-Based Decoding,” in Proc. IEEE ASRU , Scottsdale, AZ, Dec. 2015, pp. 167–174
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Sundermeyer, H. Ney, and R. Schlüter, “From Feedforward to Recurrent LSTM Neural Networks for Language Modeling,” IEEE/ACM Trans. Audio, Speech, and Language Processing , vol. 23, no. 3, pp. 517–529, Mar. 2015
2015
Earlier work this paper cites.
Y. Tachioka and S. Watanabe, “Discriminative Method for Recurrent Neural Network Language Models,” in Proc. IEEE ICASSP , South Brisbane, Australia, Apr. 2015, pp. 5386–5390
2015
Earlier work this paper cites.
J. Cui, B. Kingsbury, B. Ramabhadran, A. Sethy, K. Audhkhasi, X. Cui, E. Kislal, L. Mangu, M. Nussbaum-Thom, M. Picheny, Z. Tüske, P. Golik, R. Schlüter, H. Ney, M. J. F. Gales, K. M. Knill, A. Ragni, H. Wang, and P. Woodland, “Multilingual Representations for Low Resource Speech Recognition and Keyword Search,” in Proc. IEEE ASRU , Scottsdale, AZ, Dec. 2015, pp. 259–266
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR Corpus Based on Public Domain Audio Books,” in Proc. IEEE ICASSP , Queensland, Australia, Apr. 2015, pp. 5206–5210
2015
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,” in Proc. IEEE ICASSP , Shanghai, China, Mar. 2016, pp. 4960–4964
2016
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” in Proc. ACL , Florence, Italy, Jul. 2019, pp. 4171–4186
2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019, openAI blog. [Online]. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
2019
Later among the works it cites.
S. Kim, S. Dalmia, and F. Metze, “Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion,” in Proc. ACL , Florence, Italy, Jul. 2019, pp. 1131–1141
2019
Later among the works it cites.
A. Hannun, A. Lee, Q. Xu, and R. Collobert, “Sequence-to-Sequence Speech Recognition with Time-Depth Separable Convolutions,” in Proc. Interspeech , Graz, Austria, Sep. 2019, pp. 3785–3789
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
D. Povey, V. Peddinti, D. Galvez, P. Ghahremani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely Sequence-Trained Neural Networks for ASR Based on Lattice-Free MMI,” in Proc. Interspeech . San Francisco, CA: ISCA, Sep. 2016, pp. 2751–2755. [Online]. Available: https://doi.org/10.21437/Interspeech.2016-595
2016
Cited alongside, same era.
2016
Cited alongside, same era.
N. Jaitly, Q. V. Le, O. Vinyals, I. Sutskever, D. Sussillo, and S. Bengio, “An Online Sequence-to-Sequence Model Using Partial Conditioning,” in Proc. NIPS , Barcelona, Spain, Dec. 2016, pp. 5067–5075
2016
Cited alongside, same era.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-End Attention-Based Large Vocabulary Speech Recognition,” in Proc. IEEE ICASSP , Shanghai, China, Mar. 2016, pp. 4945–4949
2016
Cited alongside, same era.
2016
Cited alongside, same era.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, T. Han, A. Hannun, B. Jun, P. LeGresley, L. Lin, S. Narang, A. Ng, S. Ozair, R. Prenger, J. Raiman, S. Satheesh, D. Seetapun, S. Sengupta, Y. Wang, Z. Wang, C. Wang, B. Xiao, D. Yogatama, J. Zhan, and Z. Zhu, “Deep Speech 2: End-to-End Speech Recognition in English and Mandarin,” in Proc. ICML , New York City, NY, Jun. 2016, pp. 173–182
2016
Cited alongside, same era.
Y. Gal and Z. Ghahramani, “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning,” in Proc. ICML , New York City, NY, Jun. 2016, pp. 1050–1059
2016
Cited alongside, same era.
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger, “Deep Networks with Stochastic Depth,” in European Conference on Computer Vision , Amsterdam, Netherlands, Oct. 2016, pp. 646–661
2016
Cited alongside, same era.
M. Li, M. Liu, and H. Masanori, “End-to-End Speech Recognition with Adaptive Computation Steps,” in Proc. IEEE ICASSP , Brighton, UK, May 2019, pp. 6246–6250
2019
Later among the works it cites.
J. Jorge, A. Giménez, J. Iranzo-Sánchez, J. Civera, A. Sanchis, and A. Juan, “Real-Time One-Pass Decoder for Speech Recognition Using LSTM Language Models,” in Proc. Interspeech , Graz, Austria, Sep. 2019, pp. 3820–3824
2019
Later among the works it cites.
F. Stahlberg and B. Byrne, “On NMT Search Errors and Model Errors: Cat Got Your Tongue?” in Proc. EMNLP . Hong Kong, China: Association for Computational Linguistics, Nov. 2019, pp. 3354–3360
2019
Later among the works it cites.
D. Le, X. Zhang, W. Zheng, C. Fügen, G. Zweig, and M. L. Seltzer, “From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition,” in Proc. IEEE ASRU , Sentosa, Singapore, Dec. 2019, pp. 457–464
2019
Later among the works it cites.
F. Weninger, J. Andrés-Ferrer, X. Li, and P. Zhan, “Listen, Attend, Spell and Adapt: Speaker Adapted Sequence-to-Sequence ASR,” in Proc. Interspeech . Graz, Austria: ISCA, Sep. 2019, pp. 3805–3809
2019
Later among the works it cites.
Z. Meng, Y. Gaur, J. Li, and Y. Gong, “Speaker Adaptation for Attention-Based End-to-End Speech Recognition,” in Proc. Interspeech . Graz, Austria: ISCA, Sep. 2019, pp. 241–245
2019
Later among the works it cites.
H. Xu, S. Ding, and S. Watanabe, “Improving End-to-End Speech Recognition with Pronunciation-Assisted Sub-Word Modeling,” in Proc. IEEE ICASSP , Brighton, UK, Sep. 2019, pp. 7110–7114
2019
Later among the works it cites.
C. Lüscher, E. Beck, K. Irie, M. Kitza, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “RWTH ASR Systems for LibriSpeech: Hybrid vs Attention,” in Proc. Interspeech , Graz, Austria, Sep. 2019, pp. 231–235
2019
Later among the works it cites.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” in Proc. Interspeech , Graz, Austria, Sep. 2019, pp. 2613–2617
2019
Later among the works it cites.
O. Adams, M. Wiesner, S. Watanabe, and D. Yarowsky, “Massively Multilingual Adversarial Speech Recognition,” in Proc. NAACL , Minneapolis, MN, Jun. 2019, pp. 96–108
2019
Later among the works it cites.
A. Kannan, A. Datta, T. N. Sainath, E. Weinstein, B. Ramabhadran, Y. Wu, A. Bapna, Z. Chen, and S. Lee, “Large-Scale Multilingual Speech Recognition with a Streaming End-to-End Model,” in Proc. Interspeech , Graz, Austria, Sep. 2019, pp. 2130–2134
2019
Later among the works it cites.
M. Kitza, P. Golik, R. Schlüter, and H. Ney, “Cumulative Adaptation for BLSTM Acoustic Models,” in Interspeech , Graz, Austria, Sep. 2019, pp. 754–758
2019
Later among the works it cites.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A Comparison of Transformer and LSTM Encoder Decoder Models for ASR,” in Proc. IEEE ASRU , Sentosa, Singapore, Dec. 2019, pp. 8–15
2019
Later among the works it cites.
K. Kim, K. Lee, D. Gowda, J. Park, S. Kim, S. Jin, Y.-Y. Lee, J. Yeo, D. Kim, S. Jung, J. Lee, M. Han, and C. Kim, “Attention Based On-Device Streaming Speech Recognition with Large Speech Corpus,” in Proc. IEEE ASRU , Sentosa, Singapore, Dec. 2019, pp. 956–963
2019
Later among the works it cites.
T. Hori, R. Astudillo, T. Hayashi, Y. Zhang, S. Watanabe, and J. Le Roux, “Cycle-Consistency Training for End-to-End Speech Recognition,” in Proc. IEEE ICASSP , Brighton, UK, May 2019, pp. 6271–6275
2019
Later among the works it cites.
X. Chang, W. Zhang, Y. Qian, J. Le Roux, and S. Watanabe, “MIMO-Speech: End-to-End Multi-Channel Multi-Speaker Speech Recognition,” in Proc. IEEE ASRU . Sentosa, Singapore: IEEE, Dec. 2019, pp. 237–244
2019
Later among the works it cites.
“Cambridge Dictionary,” https://dictionary.cambridge.org/dictionary/english/end-to-end , accessed: 2020-02-21
2020
Later among the works it cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Proc. NeurIPS , Vancouver, BC, Canada, Dec. 2020, pp. 12 449–12 460
2020
Later among the works it cites.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid Autoregressive Transducer (HAT),” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 6139–6143
2020
Later among the works it cites.
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “A New Training Pipeline for an Improved Neural Transducer,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 2812–2816
2020
Later among the works it cites.
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, “Deliberation Model Based Two-Pass End-to-End Speech Recognition,” in Proc. IEEE ICASSP . Barcelona, Spain: IEEE, May 2020, pp. 7799–7803
2020
Later among the works it cites.
W. Han, Z. Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu, “ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 3610–3614
2020
Later among the works it cites.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7829–7833
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-Augmented Transformer for Speech Recognition,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 5036–5040
2020
Later among the works it cites.
M. Ghodsi, X. Liu, J. Apfel, R. Cabrera, and E. Weinstein, “RNN-Transducer with Stateless Prediction Network,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7049–7053
2020
Later among the works it cites.
B. Li, S.-y. Chang, T. N. Sainath, R. Pang, Y. He, T. Strohman, and Y. Wu, “Towards Fast and Accurate Streaming End-To-End ASR,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 6069–6073
2020
Later among the works it cites.
T. Yoshimura, T. Hayashi, K. Takeda, and S. Watanabe, “End-To-End Automatic Speech Recognition Integrated with CTC-Based Voice Activity Detection,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 6999–7003
2020
Later among the works it cites.
C. Weng, C. Yu, J. Cui, C. Zhang, and D. Yu, “Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 966–970
2020
Later among the works it cites.
W. Hou, Y. Dong, B. Zhuang, L. Yang, J. Shi, and T. Shinozaki, “Large-scale end-to-end multilingual speech recognition and language identification with multi-task learning,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 1037–1041
2020
Later among the works it cites.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, “Single Headed Attention Based Sequence-to-Sequence Model for State-of-the-Art Results on Switchboard,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 551–555
2020
Later among the works it cites.
W. Zhang, X. Chang, Y. Qian, and S. Watanabe, “Improving End-to-End Single-Channel Multi-Talker Speech Recognition,” IEEE/ACM Trans. Audio, Speech, and Language Processing , vol. 28, pp. 1385–1394, 2020
2020
Later among the works it cites.
C. Wang, Y. Wu, Y. Du, J. Li, S. Liu, L. Lu, S. Ren, G. Ye, S. Zhao, and M. Zhou, “Semantic Mask for Transformer Based End-to-End Speech Recognition,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 971–975
2020
Later among the works it cites.
Y. Higuchi, S. Watanabe, N. Chen, T. Ogawa, and T. Kobayashi, “Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 3655–3659
2020
Later among the works it cites.
W. Chan, C. Saharia, G. Hinton, M. Norouzi, and N. Jaitly, “Imputer: Sequence Modelling via Imputation and Dynamic Programming,” in Proc. ICML . PMLR, Jul. 2020, pp. 1403–1413
2020
Later among the works it cites.
Y. Fujita, S. Watanabe, M. Omachi, and X. Chang, “Insertion-Based Modeling for End-to-End Automatic Speech Recognition,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 3660–3664
2020
Later among the works it cites.
L. Dong and B. Xu, “Cif: Continuous Integrate-and-Fire for End-to-End Speech Recognition,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 6079–6083
2020
Later among the works it cites.
W. Zhou, R. Schlüter, and H. Ney, “Robust Beam Search for Encoder-Decoder Attention Based Speech Recognition without Length Bias,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 1768–1772
2020
Later among the works it cites.
G. Saon, Z. Tüske, and K. Audhkhasi, “Alignment-Length Synchronous Decoding for RNN Transducer,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7804–7808
2020
Later among the works it cites.
N. Moritz, T. Hori, and J. Le, “Streaming Automatic Speech Recognition with the Transformer Model,” in Proc. IEEE ICASSP . Barcelona, Spain: IEEE, May 2020, pp. 6074–6078
2020
Later among the works it cites.
L. Lu, C. Liu, J. Li, and Y. Gong, “Exploring Transformers for Large-Scale Speech Recognition,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 5041–5045
2020
Later among the works it cites.
J. Kim, Y. Lee, and E. Kim, “Accelerating RNN Transducer Inference via Adaptive Expansion Search,” IEEE Signal Processing Letters , vol. 27, pp. 2019–2023, 2020
2020
Later among the works it cites.
J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap, “Compressive Transformers for Long-Range Sequence Modelling,” Advances in Neural Information Processing Systems , vol. 33, pp. 6154–6158, 2020
2020
Later among the works it cites.
Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal Language Model Estimation for Domain-Adaptive End-to-End Speech Recognition,” in Proc. IEEE SLT , Shenzhen , China, Dec. 2020, pp. 243–250
2020
Later among the works it cites.
J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff, “Masked Language Model Scoring,” in Proc. ACL , Jul. 2020, pp. 2699–2712
2020
Later among the works it cites.
P. Bahar, N. Makarov, A. Zeyer, R. Schüter, and H. Ney, “Exploring a Zero-Order Direct HMM Based on Latent Attention for Automatic Speech Recognition,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7854–7858
2020
Later among the works it cites.
W. Zhou, R. Schlüter, and H. Ney, “Full-Sum Decoding for Hybrid HMM Based Speech Recognition Using LSTM Language Model,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7834–7838
2020
Later among the works it cites.
L. Sarı, N. Moritz, T. Hori, and J. Le Roux, “Unsupervised Speaker Adaptation using Attention-based Speaker Memory for End-to-End ASR,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 2–6
2020
Later among the works it cites.
D. Park, Y. Zhang, Y. Jia, W. Han, C.-C. Chiu, B. Li, Y. Wu, and Q. Le, “Improved Noisy Student Training for Automatic Speech Recognition,” in Proc. Interspeech , Shanghai, China, Oct. 2020, pp. 2817–2821
2020
Later among the works it cites.
W. Zhou, W. Michel, K. Irie, M. Kitza, R. Schlüter, and H. Ney, “The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7839–7843
2020
Later among the works it cites.
Y. Wang, A. Mohamed, D. Le, C. Liu, A. Xiao, J. Mahadeokar, H. Huang, A. Tjandra, X. Zhang, F. Zhang, C. Fuegen, G. Zweig, and M. L. Seltzer, “Transformer-Based Acoustic Modeling for Hybrid Speech Recognition,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 6874–6878
2020
Later among the works it cites.
J. Kahn, M. Riviere, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazare, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-Light: A Benchmark for ASR with Limited or no Supervision,” in Proc. IEEE ICASSP , Barcelona, Spain, May 2020, pp. 7669–7673
2020
Later among the works it cites.
R. Hsiao, D. Can, T. Ng, R. Travadi, and A. Ghoshal, “Online Automatic Speech Recognition with Listen, Attend and Spell Model,” IEEE Signal Processing Letters , vol. 27, pp. 1889–1893, 2020
2020
Later among the works it cites.
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, A. Bruguier, S.-y. Chang, W. Li, R. Alvarez, Z. Chen, C.-C. Chiu, D. Garcia, A. Gruenstein, K. Hu, M. Jin, A. Kannan, Q. Liang, I. McGraw, C. Peyser, R. Prabhavalkar, G. Pundak, D. Rybach, Y. Shangguan, Y. Sheth, T. Strohman, M. Visontai, Y. Wu, Y. Zhang, and D. Zhao, “A Streaming On-Device End-To-End Model Surpassing Server-Side Conventional Model Quality and Latency,” in Proc. IEEE ICASSP , Barcelona, Spain, may 2020, pp. 6059–6063
2020
Later among the works it cites.
2021
Later among the works it cites.
W. Zhou, A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “Equivalence of Segmental and Neural Transducer Modeling: A Proof of Concept,” in Proc. Interspeech , Brno, Czechia, Aug. 2021, pp. 2891–2895
2021
Later among the works it cites.
A. Narayanan, T. N. Sainath, R. Pang, J. Yu, C.-C. Chiu, R. Prabhavalkar, E. Variani, and T. Strohman, “Cascaded Encoders for Unifying Streaming and Non-Streaming ASR,” in Proc. IEEE ICASSP , Toronto, Ontario, Canada, Jun. 2021, pp. 5629–5633
2021
Later among the works it cites.
P. Guo, F. Boyer, X. Chang, T. Hayashi, Y. Higuchi, H. Inaguma, N. Kamo, C. Li, D. Garcia-Romero, J. Shi, J. Shi, S. Watanabe, K. Wei, W. Zhang, and Y. Zhang, “Recent Developments on ESPNET Toolkit Boosted by Conformer,” in Proc. IEEE ICASSP . Toronto, Ontario, Canada: IEEE, Jun. 2021, pp. 5874–5878
2021
Later among the works it cites.
R. Botros, T. Sainath, R. David, E. Guzman, W. Li, and Y. He, “Tied & Reduced RNN-T Decoder,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 4563–4567
2021
Later among the works it cites.
W. Zhou, S. Berger, R. Schlüter, and H. Ney, “Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition,” in Proc. IEEE ICASSP , Toronto, Ontario, Canada, Jun. 2021, pp. 5644–5648
2021
Later among the works it cites.
R. Prabhavalkar, Y. He, D. Rybach, S. Campbell, A. Narayanan, T. Strohman, and T. N. Sainath, “Less is More: Improved RNN-T Decoding Using Limited Label Context and Path Merging,” in Proc. IEEE ICASSP , Toronto, Ontario, Canada, Jun. 2021, pp. 5659–5663
2021
Later among the works it cites.
Y. Fujita, T. Wang, S. Watanabe, and M. Omachi, “Toward Streaming ASR with Non-Autoregressive Insertion-Based Model,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 3740–3744
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Zeineldeen, A. Glushko, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “Investigating Methods to Improve Language Model Integration for Attention-Based Encoder-Decoder ASR Models,” in Proc. Interspeech , Brno, Czechia, Aug. 2021, pp. 2856–2860
2021
Later among the works it cites.
B. Li, R. Pang, T. N. Sainath, A. Gulati, Y. Zhang, J. Qin, P. Haghani, W. R. Huang, M. Ma, and J. Bai, “Scaling End-to-End Models for Large-Scale Multilingual ASR,” in Proc. IEEE ASRU , 2021, pp. 1011–1018
2021
Later among the works it cites.
T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey, “Meta-Learning in Neural Networks: A Survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. PP, pp. 1–20, 2021
2021
Later among the works it cites.
J. Lee and S. Watanabe, “Intermediate Loss Regularization for CTC-Based Speech Recognition,” in Proc. IEEE ICASSP , Toronto, Ontario, Canada, Jun. 2021, pp. 6224–6228
2021
Later among the works it cites.
L. Meng, J. Xu, X. Tan, J. Wang, T. Qin, and B. Xu, “MixSpeech: Data Augmentation for Low-resource Automatic Speech Recognition,” in Proc. IEEE ICASSP . Toronto, Ontario, Canada: IEEE, Jun. 2021, pp. 7008–7012
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Nozaki and T. Komatsu, “Relaxing the Conditional Independence Assumption of CTC-Based ASR by Conditioning on Intermediate Predictions,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 3735–3739
2021
Later among the works it cites.
2021
Later among the works it cites.
T. Wang, Y. Fujita, X. Chang, and S. Watanabe, “Streaming End-to-End ASR Based on Blockwise Non-Autoregressive Models,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 3755–3759
2021
Later among the works it cites.
E. Tsunoo, Y. Kashiwagi, and S. Watanabe, “Streaming Transformer ASR with Blockwise Synchronous Beam Search,” in Proc. IEEE SLT . Shenzhen, China: IEEE, Jun. 2021, pp. 22–29
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Yao, D. Wu, X. Wang, B. Zhang, F. Yu, C. Yang, Z. Peng, X. Chen, L. Xie, and X. Lei, “WeNet: Production Oriented Streaming and Non-Streaming End-to-End Speech Recognition Toolkit,” Brno, Czechia, pp. 4054–4058, Sep. 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
W. Zhou, M. Zeineldeen, Z. Zheng, R. Schlüter, and H. Ney, “Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition,” in Proc. Interspeech , Brno, Czechia, Aug. 2021, pp. 2886–2890
2021
Later among the works it cites.
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Trans. Audio, Speech, and Language Processing , vol. 19, pp. 3451–3460, 2021
2021
Later among the works it cites.
E. G. Ng, C.-C. Chiu, Y. Zhang, and W. Chan, “Pushing the Limits of Non-Autoregressive Speech Recognition,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 3725–2729
2021
Later among the works it cites.
Y. Shi, Y. Wang, C. Wu, C.-F. Yeh, J. Chan, F. Zhang, D. Le, and M. Seltzer, “Emformer: Efficient Memory Transformer based Acoustic Model for Low Latency Streaming Speech Recognition,” in Proc. IEEE ICASSP . Toronto, Ontario, Canada: IEEE, Jun. 2021, pp. 6783–6787
2021
Later among the works it cites.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset,” in Proc. IEEE ICASSP . Toronto, Ontario, Canada: IEEE, Jun. 2021, pp. 5904–5908
2021
Later among the works it cites.
B. Li, A. Gulati, J. Yu, T. N. Sainath, C.-C. Chiu, A. Narayanan, S.-Y. Chang, R. Pang, Y. He, J. Qin, W. Han, Q. Liang, Y. Zhang, T. Strohman, and Y. Wu, “A Better and Faster End-to-End Model for Streaming ASR,” in Proc. IEEE ICASSP , Toronto, Ontario, Canada, Jun. 2021, pp. 5634–5638
2021
Later among the works it cites.
T. N. Sainath, Y. He, A. Narayanan, R. Botros, R. Pang, D. Rybach, C. Allauzen, E. Variani, J. Qin, Q.-N. Le-The, S.-Y. Chang, B. Li, A. Gulati, J. Yu, C.-C. Chiu, D. Caseiro, W. Li, Q. Liang, and P. Rondon, “An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 1777–1781
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Later among the works it cites.
N. Moritz, T. Hori, S. Watanabe, and J. Le Roux, “Sequence Transduction with Graph-Based Supervision,” in Proc. IEEE ICASSP , Singapore, May 2022, pp. 7212–7216
2022
Later among the works it cites.
Y. Peng, S. Dalmia, I. Lane, and S. Watanabe, “Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding,” in Proc. ICML . Baltimore, MD: PMLR, Jul. 2022, pp. 17 627–17 643
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Laptev, S. Majumdar, and B. Ginsburg, “CTC Variations Through New WFST Topologies,” in Proc. Interspeech , Incheon, Korea, sep 2022. [Online]. Available: https://doi.org/10.21437
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Yang, W. Zhou, R. Schlüter, and H. Ney, “Lattice-Free Sequence Discriminative Training for Phoneme-based Neural Transducers,” in Proc. IEEE ICASSP , Rhodes, Greece, Oct. 2022, submitted, preprint: arXiv:NNNN.NNNNN
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
A. Radford, J. W. Kim, C. McLeavey, P. Mishkin, T. Xu, G. Brockman, and I. Sutskever, “Introducing Whisper - Robust Speech Recognition via Large-Scale Weak Supervision,” Sep. 2022. [Online]. Available: https://openai.com/blog/whisper/
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Zeyer, A. Merboldt, W. Michel, R. Schlüter, and H. Ney, “Librispeech Transducer Model with Internal Language Model Prior Correction,” in Proc. Interspeech , Brno, Czech Republic, Apr. 2021, pp. 2052–2056
2056
Closest in time.
Z. Tüske, G. Saon, and B. Kingsbury, “On the Limit of English Conversational Speech Recognition,” in Proc. Interspeech , Brno, Czechia, Sep. 2021, pp. 2062–2066
2066
Closest in time.