Fetching the paper…
Reading the bibliography…
Segmental models are an alternative to frame-based models for sequence prediction, where hypothesized path weights are based on entire segment scores rather than a single frame at a time.
F. Jelinek, “Continuous speech recognition by statistical methods,” Proceedings of the IEEE , vol. 64, no. 4, pp. 532–556, 1976
1976
Earlier work this paper cites.
L. Bahl, F. Jelinek, and R. Mercer, “A maximum likelihood approach to speech recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 5, pp. 179–190, 1983
1983
Earlier work this paper cites.
M. A. Bush and G. E. Kopec, “Network-based connected digit recognition using vector quantization,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1985
1985
Earlier work this paper cites.
K.-F. Lee, “On large-vocabulary speaker-independent continuous speech recognition,” Speech communication , vol. 7, no. 4, pp. 375–379, 1988
1988
Earlier work this paper cites.
L. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,” Proceedings of the IEEE , vol. 77, no. 2, pp. 257–286, 1989
1989
Earlier work this paper cites.
V. Zue, J. Glass, M. Phillips, and S. Seneff, “Acoustic segmentation and phonetic classification in the SUMMIT system,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 1989
1989
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,” NASA STI/Recon technical report , vol. 93, 1993
1993
Earlier work this paper cites.
M. Ostendorf, V. Digalakis, and O. Kimball, “From HMM’s to segment models: A unified view of stochastic modeling for speech recognition,” IEEE Transactions on Speech and Audio Processing , pp. 360–378, 1996
1996
Earlier work this paper cites.
G. Chung and S. Seneff, “Hierarchical duration modelling for speech recognition using the ANGIE framework.” in Eurospeech , 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
M. Mohri, “Finite-state transducers in language and speech processing,” Computational linguistics , vol. 23, no. 2, pp. 269–311, 1997
1997
Earlier work this paper cites.
A. Halberstadt and J. Glass, “Heterogeneous measurements and multiple classifiers for speech recognition,” in International Conference on Spoken Language Processing , 1998
1998
Earlier work this paper cites.
N. Smith and M. Gales, “Speech recognition using SVMs,” in Advances in neural information processing systems (NIPS) , 2001
2001
Earlier work this paper cites.
——, “Semiring frameworks and algorithms for shortest-distance problems,” Journal of Automata, Languages and Combinatorics , vol. 7, no. 3, pp. 321–350, 2002
2002
Earlier work this paper cites.
J. R. Glass, “A probabilistic framework for segment-based speech recognition,” Computer Speech & Language , vol. 17, no. 2, pp. 137–152, 2003
2003
Earlier work this paper cites.
S. Sarawagi and W. W. Cohen, “Semi-Markov conditional random fields for information extraction.” in Advances in Neural Information Processing Systems (NIPS) , vol. 17, 2004
2004
Earlier work this paper cites.
M. Hasegawa-Johnson, J. Baker, S. Borys, K. Chen, E. Coogan, S. Greenberg, A. Juneja, K. Kirchhoff, K. Livescu, S. Mohan et al. , “Landmark-based speech recognition: Report of the 2004 Johns Hopkins summer workshop,” in IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2005
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in International Conference on Machine Learning (ICML) , 2006
2006
Earlier work this paper cites.
M. De Wachter, M. Matton, K. Demuynck, P. Wambacq, R. Cools, and D. Van Compernolle, “Template-based continuous speech recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 4, pp. 1377–1390, 2007
2007
Earlier work this paper cites.
G. Heigold, T. Deselaers, R. Schlüter, and H. Ney, “Modified MMI/MPE: A direct evaluation of the margin in speech recognition,” in International Conference on Machine learning (ICML) , 2008
2008
Cited alongside, same era.
G. Zweig and P. Nguyen, “A segmental CRF approach to large vocabulary continuous speech recognition,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2009
2009
Cited alongside, same era.
E. McDermott, S. Watanabe, and A. Nakamura, “Discriminative training based on an integrated view of MPE and MMI in margin and error space,” in IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) , 2010
2010
Cited alongside, same era.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in AISTATS , 2010
2010
Cited alongside, same era.
Y. Miao, J. Li, Y. Wang, S. Zhang, and Y. Gong, “Simplifying long short-term memory acoustic models for fast training and decoding,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015
2015
Later among the works it cites.
Y. He, “Segmental models with an exploration of acoustic and lexical grouping in automatic speech recognition,” Ph.D. dissertation, The Ohio State University, 2015
2015
Later among the works it cites.
L. Tóth, “Phone recognition with hierarchical convolutional deep maxout networks,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 1, p. 25, 2015
2015
Later among the works it cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The Kaldi speech recognition toolkit,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2011
2011
Cited alongside, same era.
G. Zweig, “Classification and recognition with direct segment models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2012
2012
Cited alongside, same era.
Y. He and E. Fosler-Lussier, “Efficient segmental conditional random fields for phone recognition,” in INTERSPEECH , 2012
2012
Cited alongside, same era.
T. Tieleman and G. Hinton, “Lecture 6.5-RMSprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural networks for machine learning , 2012
2012
Cited alongside, same era.
A. Graves, “Sequence transduction with recurrent neural networks,” CoRR , vol. abs/1211.3711, 2012
2012
Cited alongside, same era.
A. L. Maas, S. D. Miller, T. M. O’neil, A. Y. Ng, and P. Nguyen, “Word-level acoustic modeling with convolutional vector regression,” in ICML Workshop on Representation Learning , 2012
2012
Cited alongside, same era.
E. Fosler-Lussier, Y. He, P. Jyothi, and R. Prabhavalkar, “Conditional random fields in speech, audio, and language processing,” Proceedings of the IEEE , vol. 101, no. 5, pp. 1054–1075, 2013
2013
Cited alongside, same era.
O. Abdel-Hamid, L. Deng, D. Yu, and H. Jiang, “Deep segmental neural networks for speech recognition,” in INTERSPEECH , 2013
2013
Cited alongside, same era.
Y. He and E. Fosler-Lussier, “Segmental conditional random fields with deep neural networks as acoustic models for first-pass word recognition.” in INTERSPEECH , 2015
2015
Later among the works it cites.
A. L. Maas, Z. Xie, D. Jurafsky, and A. Y. Ng, “Lexicon-free conversational speech recognition with neural networks,” in Human Language Technologies: The Annual Conference of the North American Chapter of the ACL , 2015
2015
Later among the works it cites.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015
2015
Later among the works it cites.
L. Lu, L. Kong, C. Dyer, N. A. Smith, and S. Renals, “Segmental recurrent neural networks for end-to-end speech recognition,” in INTERSPEECH , 2016
2016
Later among the works it cites.
L. Kong, C. Dyer, and N. A. Smith, “Segmental recurrent neural networks,” in International Conference on Learning Representations (ICLR) , 2016
2016
Later among the works it cites.
H. Tang, W. Wang, K. Gimpel, and K. Livescu, “End-to-end training approaches for discriminative segmental models,” in IEEE Workshop on Spoken Language Technology (SLT) , 2016
2016
Later among the works it cites.
H. Tang, W. Wang, K. Gimpel, and K. Livescu, “Efficient segmental cascades for speech recognition,” in INTERSPEECH , 2016
2016
Later among the works it cites.
P. Doetsch, S. Hegselmann, R. Schlüter, and H. Ney, “Inverted HMM – a proof of concept,” in NIPS 2016 End-to-end Learning for Speech and Audio Processing Workshop , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Later among the works it cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2016
2016
Later among the works it cites.
D. Povey, V. Peddinti, D. Galvez, P. Ghahrmani, V. Manohar, X. Na, Y. Wang, and S. Khudanpur, “Purely sequence-trained neural networks for ASR based on lattice-free MMI,” in INTERSPEECH , 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
2017
Closest in time.