Fetching the paper…
Reading the bibliography…
Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks.
Neural Network Adaptive Beamforming for Robust Multichannel Speech Recognition
Li, B.; Sainath, T.; et al. 2016 · 1980
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; et al. 2006 · 2006
Earlier work this paper cites.
Dlib-ml: A machine learning toolkit
King, D. E. 2009 · 2009
Earlier work this paper cites.
Microphone array processing for distant speech recognition: From close-talking microphones to far-field sensors
Kumatani, K.; McDonough, J.; et al. 2012 · 2012
Earlier work this paper cites.
An overview of noise-robust automatic speech recognition
Li, J.; Deng, L.; et al. 2014 · 2014
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition
Bahdanau, D.; Chorowski, J.; et al. 2016 · 2016
Earlier work this paper cites.
Neural network based spectral mask estimation for acoustic beamforming
Heymann, J.; Drude, L.; et al. 2016 · 2016
Earlier work this paper cites.
Deep beamforming networks for multi-channel speech recognition
Xiao, X.; Watanabe, S.; et al. 2016 · 2016
Earlier work this paper cites.
Multichannel end-to-end speech recognition
Ochiai, T.; Watanabe, S.; et al. 2017 · 2017
Earlier work this paper cites.
NARA-WPE: A Python package for weighted prediction error dereverberation in Numpy and Tensorflow for online and offline processing
Drude, L.; Heymann, J.; et al. 2018 · 2018
Earlier work this paper cites.
3-D CNN models for far-field multi-channel speech recognition
Ganapathy, S.; and Peddinti, V. 2018 · 2018
Earlier work this paper cites.
Multi-geometry spatial acoustic modeling for distant speech recognition
Kenichi, K.; Wu, M.; et al. 2019 · 2019
Cited alongside, same era.
An iterative mask estimation approach to deep learning based multi-channel speech recognition
Tu, Y.; Du, J.; et al. 2019 · 2019
Cited alongside, same era.
Frequency domain multi-channel acoustic modeling for distant speech recognition
Wu, M.; Kenichi, K.; et al. 2019 · 2019
Cited alongside, same era.
Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Baevski, A.; Zhou, Y.; et al. 2020 · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
Gulati, A.; Qin, J.; et al. 2020 · 2020
Cited alongside, same era.
Lipreading Using Temporal Convolutional Networks
Martinez, B.; Ma, P.; et al. 2020 · 2020
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Chen, S.; Wang, C.; et al. 2022 · 2022
Later among the works it cites.
End-to-End Audio-Visual Neural Speaker Diarization
He, M.; Du, J.; et al. 2022 · 2022
Later among the works it cites.
Visual speech recognition for multiple languages in the wild
Ma, P.; Petridis, S.; et al. 2022 · 2022
Later among the works it cites.
The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition
Wang, Z.; Wu, S.; et al. 2023 · 2022
Later among the works it cites.
WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition
Zhang, B.; Lv, H.; et al. 2022 · 2022
Later among the works it cites.
SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robust multi-channel speech recognition using frequency aligned network
Park, T.; Kumatani, K.; et al. 2020 · 2020
Cited alongside, same era.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Hsu, W.; Bolte, B.; et al. 2021 · 2021
Cited alongside, same era.
End-To-End Audio-Visual Speech Recognition with Conformers
Ma, P.; Petridis, S.; et al. 2021 · 2021
Cited alongside, same era.
Acoustic Beamforming for Speaker Diarization of Meetings
Anguera, X.; Wooters, C.; and Hernando, J. 2007 · 2022
Cited alongside, same era.
Audio-Visual Speech Recognition in MISP2021 Challenge: Dataset Release and Deep Analysis
Chen, H.; Du, J.; et al. 2022 · 2022
Cited alongside, same era.
The First Multimodal Information Based Speech Processing (Misp) Challenge: Data, Tasks, Baselines And Results
Chen, H.; Zhou, H.; et al. 2022 · 2022
Cited alongside, same era.
Zhang, Z.; Zhou, L.; et al. 2022 · 2022
Later among the works it cites.
MIR-GAN: Refining Frame-Level Modality-Invariant Representations with Adversarial Network for Audio-Visual Speech Recognition
Hu, Y.; Chen, C.; et al. 2023 · 2023
Later among the works it cites.
Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition
Hu, Y.; Li, R.; Chen, C.; et al. 2023 · 2023
Later among the works it cites.
Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
Hu, Y.; Li, R.; et al. 2023 · 2023
Later among the works it cites.
Lian, J.; Baevski, A.; et al. 2023 · 2023
Later among the works it cites.
VatLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
Zhu, Q.; Zhou, L.; et al. 2023 · 2023
Later among the works it cites.