Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks.
“A data-driven organization of the dynamic programming beam search for continuous speech recognition,”
Hermann Ney et al., · 1987
Earlier work this paper cites.
“Speaker diarization from speech transcripts,”
Leonardo Canseco-Rodriguez, Lori Lamel, and Jean-Luc Gauvain, · 2004
Earlier work this paper cites.
“A comparative study using manual and automatic transcriptions for diarization,”
Leonardo Canseco, Lori Lamel, and J-L Gauvain, · 2005
Earlier work this paper cites.
“The AMI meeting corpus,”
Wessel Kraaij et al., · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves et al., · 2006
Earlier work this paper cites.
“KenLM: Faster and smaller language model queries,”
Kenneth Heafield et al., · 2011
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani et al., · 2017
Earlier work this paper cites.
“Word beam search: A connectionist temporal classification decoding algorithm,”
Harald Scheidl et al., · 2018
Earlier work this paper cites.
“Multimodal speaker segmentation and diarization using lexical and acoustic cues via sequence to sequence neural networks,”
Tae Jin Park and Panayiotis Georgiou, · 2018
Earlier work this paper cites.
“Joint speech recognition and speaker diarization via sequence transduction,”
Laurent El Shafey et al., · 2019
Earlier work this paper cites.
“Speaker diarization with lexical information,”
Tae Jin Park et al., · 2019
Earlier work this paper cites.
“Megatron-LM: Training multi-billion parameter language models using model parallelism,”
Mohammad Shoeybi et al., · 2019
Earlier work this paper cites.
“NeMo: a toolkit for building ai applications using neural modules,”
Oleksii Kuchaiev et al., · 2019
Cited alongside, same era.
“Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap,”
Tae Jin Park et al., · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford et al., · 2019
Cited alongside, same era.
“Optuna: A next-generation hyperparameter optimization framework,”
Takuya Akiba et al., · 2019
Cited alongside, same era.
“Target-speaker voice activity detection: a novel approach for multi-speaker diarization in a dinner party scenario,”
Ivan Medennikov et al., · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
“TitaNet: Neural model for speaker representation with 1d depth-wise separable convolutions and global context,”
Nithin Rao Koluguri et al., · 2022
Later among the works it cites.
“Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks,”
Federico Landini et al., · 2022
Later among the works it cites.
“The CHiME-7 DASR Challenge: Distant meeting transcription with multiple devices in diverse scenarios,”
Samuele Cornell et al., · 2023
Closest in time.
Luyao Cheng et al., · 2023
Closest in time.
“Encoder-decoder multimodal speaker change detection,”
Jee-weon Jung et al., · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anmol Gulati et al., · 2020
Cited alongside, same era.
“CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,”
Shinji Watanabe et al., · 2020
Cited alongside, same era.
“What’s in the box? a preliminary analysis of undesirable content in the common crawl corpus,”
Alexandra Sasha Luccioni and Joseph D Viviano, · 2021
Cited alongside, same era.
“A review of speaker diarization: Recent advances with deep learning,”
Tae Jin Park et al., · 2022
Cited alongside, same era.
“Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed asr,”
Naoyuki Kanda et al., · 2022
Cited alongside, same era.
“Turn-to-diarize: Online speaker diarization constrained by transformer transducer speaker turn detection,”
Wei Xia et al., · 2022
Cited alongside, same era.
“Asr-aware end-to-end neural diarization,”
Aparna Khare et al., · 2022
Cited alongside, same era.
Rohit Paturi et al., · 2023
Closest in time.
“The CHiME-7 Challenge: System description and performance of nemo team’s DASR system,”
Tae Jin Park et al., · 2023
Closest in time.
“PyCTCDecode: A fast and lightweight python-based CTC beam search decoder for speech recognition,” 2023,
Kensho Technologies, · 2023
Closest in time.
“LibriSpeech: an asr corpus based on public domain audio books,” 2015,
Vassil Panayotov et al., · 2023
Closest in time.
“NeMo Megatron Launcher,” https://github.com/NVIDIA/NeMo-Megatron-Launcher , 2023,
NVIDIA, · 2023
Closest in time.
“The-Stack-Dedup:Datasets at Hugging Face,” https://huggingface.co/datasets/bigcode/the-stack-dedup , 2023,
Hugging Face, · 2023
Closest in time.
“Whisper: Robust speech recognition via large-scale weak supervision,” https://github.com/openai/whisper , 2023,
OpenAI, · 2023
Closest in time.