Fetching the paper…
Reading the bibliography…
SpeechBrain is an open-source and all-in-one speech toolkit.
Lingvo: A modular and scalable framework for sequence-to-sequence modeling
J. Shen, P. Nguyen, Y. Wu, Z. Chen, M. X. Chen, Y. Jia, A. Kannan, T. Sainath, Y. Cao, C. Chiu, et al · 1902
Earlier work this paper cites.
Fairseq: A fast, extensible toolkit for sequence modeling, 2019
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 1904
Earlier work this paper cites.
NeMo: a toolkit for building AI applications using Neural Modules, 2019
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen · 1909
Earlier work this paper cites.
Y. Luo, Z. Chen, and T. Yoshioka · 1910
Earlier work this paper cites.
The generalized correlation method for estimation of time delay
C. H. Knapp and G. C. Carter · 1976
Earlier work this paper cites.
Multiple emitter location and signal parameter estimation
R. Schmidt · 1986
Earlier work this paper cites.
The development of the time-delay neural network architecture for speech recognition
K. J. Lang and G. E. Hinton · 1988
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. Lang · 1989
Earlier work this paper cites.
DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM, 1993
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren · 1993
Earlier work this paper cites.
Acoustic event localization using a crosspower-spectrum phase based technique
M. Omologo and P. Svaizer · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Julius: An open source realtime large vocabulary recognition engine
A. Lee, T. Kawahara, and K. Shikano · 2001
Earlier work this paper cites.
Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs
A.W. Rix, J.G. Beerends, M.P. Hollier, and A.P. Hekstra · 2001
Earlier work this paper cites.
The HTK Book
S. Young, G. Evermann, T. Hain, D. Kershaw, G. Moore, J. Odell, D. Ollason, D. Povey, V. Valtchev, and P. Woodland · 2002
Earlier work this paper cites.
Weighted finite-state transducers in speech recognition
M. Mohri, F. Pereira, and M. Riley · 2002
Earlier work this paper cites.
Librimix: An open-source dataset for generalizable speech separation, 2020
J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, and E. Vincent · 2005
Earlier work this paper cites.
W. Han, Z. Zhang, Y. Zhang, J. Yu, C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu · 2005
Earlier work this paper cites.
Pocketsphinx: A free, real-time continuous speech recognition system for hand-held devices
D. Huggins-Daines, M. Kumar, A. Chan, A. W. Black, M. Ravishankar, and A. I. Rudnicky · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska, I. McCowan, W. Post, D. Reidsma, and P. Wellner · 2006
Earlier work this paper cites.
Design fragments make using frameworks easier
G. Fairbanks, D. Garlan, and W. Scherlis · 2006
Earlier work this paper cites.
Performance measurement in blind audio source separation
E. Vincent, R. Gribonval, and C. Fevotte · 2006
Earlier work this paper cites.
A tutorial on spectral clustering
U. Luxburg · 2007
Earlier work this paper cites.
M. Pal, M. Kumar, R. Peri, T. Park, S. Kim, C. Lord, S. Bishop, and S. Narayanan · 2007
Earlier work this paper cites.
Evaluation of objective quality measures for speech enhancement
Y. Hu and P. Loizou · 2007
Earlier work this paper cites.
New insights into the MVDR beamformer in room acoustics
E. Habets, J. Benesty, I. Cohen, S. Gannot, and J. Dmochowski · 2009
Earlier work this paper cites.
A modified SRP-PHAT functional for robust real-time sound source localization with scalable spatial sampling
M. Cobos, A. Marti, and J. Lopez · 2010
Earlier work this paper cites.
RASR - The RWTH Aachen University Open Source Speech Recognition Toolkit
D. Rybach, S. Hahn, P. Lehnen, D. Nolden, M. Sundermeyer, Z. Tüske, S. Wiesler, R. Schlüter, and H. Ney · 2011
Earlier work this paper cites.
The Kaldi Speech Recognition Toolkit
D. Povey et al · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, et al · 2011
Earlier work this paper cites.
Analysis of i-vector length normalization in speaker recognition systems
D. Garcia-Romero and C. Espy-Wilson · 2011
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
A. Graves · 2012
Earlier work this paper cites.
F. Landini, J. Profant, M. Diez, and L. Burget · 2012
Earlier work this paper cites.
The voice bank corpus: Design, collection and data analysis of a large regional accent speech database
C. Veaux, J. Yamagishi, and S. King · 2013
Earlier work this paper cites.
PLDA for speaker verification with utterances of arbitrary duration
P. Kenny, T. Stafylakis, P. Ouellet, Md. J. Alam, and P. Dumouchel · 2013
Earlier work this paper cites.
The Diverse Environments Multi-channel Acoustic Noise Database (DEMAND): A database of multichannel environmental noise recordings
J. Thiemann, N. Ito, and E. Vincent · 2013
Earlier work this paper cites.
Deep speech: Scaling up end-to-end speech recognition, 2014
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y. Ng · 2014
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
End-to-end speech recognition with recurrent neural networks
A. Graves and N. Jaitly · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Robust speech recognition with speech enhanced deep neural networks
J. Du, Q. Wang, T. Gao, Y. Xu, L. Dai, and C. Lee · 2014
Cited alongside, same era.
Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks
Z. Chen, S. Watanabe, H. Erdogan, and J.R. Hershey · 2015
Cited alongside, same era.
Joint training of front-end and back-end deep neural networks for robust speech recognition
T. Gao, J. Du, L. Dai, and C. Lee · 2015
Cited alongside, same era.
LibriSpeech: An ASR corpus based on public domain audio books
Espresso: A fast end-to-end neural speech recognition toolkit
Y. Wang, T. Chen, H. Xu, S. Ding, H. Lv, Y. Shao, N. Peng, L. Xie, S. Watanabe, and S. Khudanpur · 2019
Later among the works it cites.
Conv-tasnet: Surpassing ideal time-frequency magnitude masking for speech separation
Y. Luo and N. Mesgarani · 2019
Later among the works it cites.
WHAM!: extending speech separation to noisy environments
G. Wichern, J. Antognini, M. Flynn, L. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. Le Roux · 2019
Later among the works it cites.
Speech model pre-training for end-to-end spoken language understanding
L. Lugosch, M. Ravanelli, P. Ignoto, V. Tomar, and Y. Bengio · 2019
Later among the works it cites.
Jasper: An end-to-end convolutional neural acoustic model
J. Li, V. Lavrukhin, B. Ginsburg, R. Leary, O. Kuchaiev, J. M. Cohen, H. Nguyen, and R. Teja Gadde · 2019
Later among the works it cites.
Pytorch Lightning
W. Falcon et al · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
An extensible speaker identification sidekit in python
A. Larcher, K. A. Lee, and S. Meignier · 2016
Cited alongside, same era.
A joint training framework for robust automatic speech recognition
Z. Wang and D. Wang · 2016
Cited alongside, same era.
Batch-normalized joint training for DNN-based distant speech recognition
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio · 2016
Cited alongside, same era.
Deep clustering: Discriminative embeddings for segmentation and separation
J. Hershey, Z. Chen, J. Le Roux, and S. Watanabe · 2016
Cited alongside, same era.
Later among the works it cites.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le · 2019
Later among the works it cites.
Libri-light: A benchmark for ASR with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2019
Later among the works it cites.
A comparative study on transformer vs rnn in speech applications
S. Karita, N. Chen, T. Hayashi, T. Hori, H. Inaguma, Z. Jiang, M. Someki, N. Soplin, R. Yamamoto, X. Wang, S. Watanabe, T. Yoshimura, and W. Zhang · 2019
Later among the works it cites.
But system description to voxceleb speaker recognition challenge 2019
H. Zeinali, S. Wang, A. Silnova, P. Matějka, and O. Plchot · 2019
Later among the works it cites.
Quaternion recurrent neural networks
T. Parcollet, M. Ravanelli, M. Morchid, G. Linarès, C. Trabelsi, R. De Mori, and Y. Bengio · 2019
Later among the works it cites.
Lightweight and optimized sound source localization and tracking methods for open and closed microphone array configurations
F. Grondin and F. Michaud · 2019
Later among the works it cites.
Asteroid: the PyTorch-based audio source separation toolkit for researchers
M. Pariente, S. Cornell, J. Cosentino, S. Sivasankaran, E. Tzinis, J. Heitkaemper, M. Olvera, F. Stöter, M. Hu, J. M. Martín-Doñas, D. Ditter, A. Frank, A. Deleforge, and E. Vincent · 2020
Later among the works it cites.
pyannote.audio: neural building blocks for speaker diarization
H. Bredin, R. Yin, J. M. Coria, G. Gelly, P. Korshunov, M. Lavechin, D. Fustes, H. Titeux, W. Bouaziz, and M. Gill · 2020
Later among the works it cites.
Common Voice: a massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber · 2020
Later among the works it cites.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
B. Desplanques, J. Thienpondt, and K. Demuynck · 2020
Later among the works it cites.
The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results
C. Reddy, V. Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun, P. Rana, S. Srinivasan, and J. Gehrke · 2020
Later among the works it cites.
Whamr!: Noisy and reverberant single-channel speech separation
M. Maciejewski, G. Wichern, E. McQuinn, and J. Le Roux · 2020
Later among the works it cites.
SLURP: A Spoken Language Understanding Resource Package
E. Bastianelli, A. Vanzo, P. Swietojanski, and V. Rieser · 2020
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Later among the works it cites.
Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush · 2020
Later among the works it cites.
SciPy 1.0: fundamental algorithms for scientific computing in Python
P. Virtanen, R. Gommers, T. E Oliphant, M. Haberland, T. Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al · 2020
Later among the works it cites.
Fastai: A layered api for deep learning
J. Howard and S. Gugger · 2020
Later among the works it cites.
WER we are and WER we think we are
P. Szymański, P. Żelasko, M. Morzy, A. Szymczak, M. Żyła-Hoppe, J. Banaszczak, L. Augustyniak, J. Mizgajski, and Y. Carmiel · 2020
Later among the works it cites.
Rethinking evaluation in ASR: are our models robust enough?
T. Likhomanenko, Q. Xu, V. Pratap, P. Tomasello, J. Kahn, G. Avidov, R. Collobert, and G. Synnaeve · 2020
Later among the works it cites.
Real time speech enhancement in the waveform domain
A. Défossez, G. Synnaeve, and Y. Adi · 2020
Later among the works it cites.
GEV beamforming supported by DOA-based masks generated on pairs of microphones
F. Grondin, J. Lauzon, J. Vincent, and F. Michaud · 2020
Later among the works it cites.
Superb: Speech processing universal performance benchmark, 2021
S. Yang, P. Chi, Y. Chuang, C. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G. Lin, T. Huang, W. Tseng, K. Lee, D. Liu, Z. Huang, S. Dong, S. Li, S. Watanabe, A. Mohamed, and H. Lee · 2021
Closest in time.
ECAPA-TDNN embeddings for speaker diarization, 2021
N. Dawalatabad, M. Ravanelli, F. Grondin, J. Thienpondt, B. Desplanques, and H. Na · 2021
Closest in time.
MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
S. Fu, C. Yu, T. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao · 2021
Closest in time.
Attention is all you need in speech separation
C. Subakan, M. Ravanelli, S. Cornell, M. Bronzi, and J. Zhong · 2021
Closest in time.
Timers and Such: A practical benchmark for spoken language understanding with numbers
L. Lugosch, P. Papreja, M. Ravanelli, A. Heba, and T. Parcollet · 2021
Closest in time.
S. Majumdar, J. Balam, O. Hrinchuk, V. Lavrukhin, V. Noroozi, and B. Ginsburg · 2021
Closest in time.
C. Wang, M. Rivière, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux · 2021
Closest in time.
Espnet-se: End-to-end speech enhancement and separation toolkit designed for ASR integration
C. Li, J. Shi, W. Zhang, A. Subramanian, X. Chang, N. Kamo, M. Hira, T. Hayashi, C. Böddeker, Z. Chen, and Shinji Watanabe · 2021
Closest in time.
S. Seo, D. Kwak, and B. Lee · 2021
Closest in time.