Fetching the paper…
Reading the bibliography…
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more.
Diffwave: A Versatile Diffusion Model for Audio Synthesis
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro · 2009
Earlier work this paper cites.
KenLM: Faster and Smaller Language Model Queries
K. Heafield · 2011
Earlier work this paper cites.
Mne software for processing meg and eeg data
A. Gramfort, M. Luessi, E. Larson, D. A. Engemann, D. Strohmeier, C. Brodbeck, L. Parkkonen, and M. S. Hämäläinen · 2014
Earlier work this paper cites.
On the selection of the impulse responses for distant-speech recognition based on contaminated speech training
M. Ravanelli and M. Omologo · 2014
Earlier work this paper cites.
Contaminated speech training methods for robust DNN-HMM distant speech recognition
M. Ravanelli and M. Omologo · 2015
Earlier work this paper cites.
Joint ctc-attention based end-to-end speech recognition using multi-task learning
S. Kim, T. Hori, and S. Watanabe · 2017
Earlier work this paper cites.
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu · 2017
Earlier work this paper cites.
EEGNet: a compact convolutional neural network for EEG-based brain computer interfaces
V. J. Lawhern, A. J. Solon, N. R. Waytowich, S. M. Gordon, C. P. Hung, and B. J. Lance · 2018
Earlier work this paper cites.
Light gated recurrent units for speech recognition
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio · 2018
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai · 2018
Earlier work this paper cites.
RNN-T for latency controlled ASR with improved beam search
M. Jain, K. Schubert, J. Mahadeokar, C. Yeh, K. Kalgaonkar, A. Sriram, C. Fuegen, and M. L. Seltzer · 2019
Earlier work this paper cites.
NeMo: a toolkit for building AI applications using Neural Modules
O. Kuchaiev, J. Li, H. Nguyen, O. Hrinchuk, R. Leary, B. Ginsburg, S. Kriman, S. Beliaev, V. Lavrukhin, J. Cook, P. Castonguay, M. Popova, J. Huang, and J. M. Cohen · 2019
Earlier work this paper cites.
Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation
Y. Luo and N. Mesgarani · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff · 2019
Earlier work this paper cites.
ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification
B. Desplanques, J. Thienpondt, and K. Demuynck · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Cited alongside, same era.
Dual-path RNN: efficient long sequence modeling for time-domain single-channel speech separation
Y. Luo, Z. Chen, and T. Yoshioka · 2020
Cited alongside, same era.
Conversational AI: Dialogue Systems, Conversational Agents, and Chatbots
M. McTear · 2021
Cited alongside, same era.
SpeechBrain: A general-purpose speech toolkit
M. Ravanelli, T. Parcollet, P. Plantinga, A. Rouhe, S. Cornell, L. Lugosch, C. Subakan, N. Dawalatabad, A. Heba, J. Zhong, et al · 2021
Cited alongside, same era.
Leakage and the reproducibility crisis in machine-learning-based science
S. Kapoor and A. Narayanan · 2023
Later among the works it cites.
Hyperconformer: Multi-head hypermixer for efficient speech recognition
F. Mai, J. Zuluaga-Gomez, T. Parcollet, and P. Motlicek · 2023
Later among the works it cites.
Stabilising and accelerating light gated recurrent units for automatic speech recognition
A. Moumen and T. Parcollet · 2023
Later among the works it cites.
Posthoc Interpretation via Quantization
F. Paissan, C. Subakan, and M. Ravanelli · 2023
Later among the works it cites.
EEG conformer: Convolutional transformer for EEG decoding and visualization
Y. Song, Q. Zheng, B. Liu, and X. Gao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fastspeech 2: Fast and high-quality end-to-end text to speech
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu · 2021
Cited alongside, same era.
Superb: Speech processing universal performance benchmark
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee · 2021
Cited alongside, same era.
Skim: Skipping memory lstm for low-latency real-time continuous speech separation
C. Li, L. Yang, W. Wang, and Y. Qian · 2022
Cited alongside, same era.
Listen to Interpret: Post-hoc Interpretability for Audio Networks with NMF
J. Parekh, S. Parekh, P. Mozharovskyi, F. Alche-Buc, and G. Richard · 2022
Cited alongside, same era.
A SpeechBrain for Everything: State of the PyTorch Ecosystem for Speech Technologies
A. Rouhe, M. Ravanelli, T. Parcollet, and P. Plantinga · 2022
Cited alongside, same era.
pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe
H. Bredin · 2023
Cited alongside, same era.
CL-MASR: A Continual Learning Benchmark for Multilingual ASR
L. Della Libera, P. Mousavi, S. Zaiem, C. Subakan, and M. Ravanelli · 2023
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
Speech Emotion Diarization: Which Emotion Appears When?
Y. Wang, M. Ravanelli, and A. Yacoubi · 2023
Later among the works it cites.
Mother of all BCI Benchmarks, 2024
B. Aristimunha, I. Carrara, P. Guetschel, S. Sedlar, P. Rodrigues, J. Sosulski, D. Narayanan, E. Bjareholt, Q. Barthelemy, R. Kobler, R. T. Schirrmeister, E. Kalunga, L. Darmet, C. Gregoire, A. Abdul Hussain, R. Gatti, V. Goncharenko, J. Thielen, T. Moreau, Y. Roy, V. Jayaram, A. Barachant, and S. Chevallier · 2024
Closest in time.
SpeechBrain-MOABB: An open-source Python library for benchmarking deep neural networks applied to EEG signals
D. Borra, F. Paissan, and M. Ravanelli · 2024
Closest in time.
Focal modulation networks for interpretable sound classification
L. Della Libera, C. Subakan, and M. Ravanelli · 2024
Closest in time.
Rethinking open source generative AI: open washing and the EU AI Act
A. Liesenfeld and M. Dingemanse · 2024
Closest in time.
Listenable Maps for Audio Classifiers
F. Paissan, M. Ravanelli, and C. Subakan · 2024
Closest in time.
LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech
T. Parcollet, H. Nguyen, S. Evain, M. Zanon Boito, A. Pupier, S. Mdhaffar, H. Le, S. Alisamir, N. Tomashenko, M. Dinarelli, S. Zhang, A. Allauzen, M. Coavoux, Y. Estève, M. Rouvier, J. Goulian, B. Lecouteux, F. Portet, S. Rossato, F. Ringeval, D. Schwab, and L. Besacier · 2024
Closest in time.
Progres: Prompted generative rescoring on asr n-best
A. D. Tur, A. Moumen, and M. Ravanelli · 2024
Closest in time.
Open Implementation and Study of BEST-RQ for Speech Processing
R. Whetten, T. Parcollet, M. Dinarelli, and Y. Estève · 2024
Closest in time.