Fetching the paper…
Reading the bibliography…
Recent studies have highlighted the importance of fully open foundation models.
“SWITCHBOARD: telephone speech corpus for research and development”
J.J. Godfrey · 1992
Earlier work this paper cites.
“The design for the Wall Street Journal-based CSR corpus”
Douglas Paul and Janet Baker · 1992
Earlier work this paper cites.
“Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus”
Jean Carletta · 2007
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov · 2015
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline”
Hui Bu · 2017
Earlier work this paper cites.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Earlier work this paper cites.
“TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation”
François Hernandez · 2018
Earlier work this paper cites.
“Common voice: A massively-multilingual speech corpus”
Rosana Ardila · 2019
Earlier work this paper cites.
“CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit”, 2019
Junichi Yamagishi · 2019
Earlier work this paper cites.
“Pytorch: An imperative style, high-performance deep learning library”
A. Paszke · 2019
Earlier work this paper cites.
“Conformer: Convolution-augmented Transformer for Speech Recognition”
Anmol Gulati et al · 2020
Earlier work this paper cites.
“Ksponspeech: Korean spontaneous speech corpus for automatic speech recognition”
Jeong-Uk Bang · 2020
Earlier work this paper cites.
“MLS: A large-scale multilingual dataset for speech research”
Vineel Pratap · 2020
Earlier work this paper cites.
“Recent developments on espnet toolkit boosted by conformer”
Pengcheng Guo et al · 2021
Earlier work this paper cites.
“VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation”
Changhan Wang · 2021
Earlier work this paper cites.
“CoVoST 2 and Massively Multilingual Speech Translation”
Changhan Wang · 2021
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark”
Shu-wen Yang et al · 2021
Cited alongside, same era.
“Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion”
Duc Le et al · 2021
Cited alongside, same era.
“PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition”
Cheng-I Lai et al · 2021
Cited alongside, same era.
“Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding”
Yifan Peng, Siddharth Dalmia, Ian Lane and Shinji Watanabe · 2022
Cited alongside, same era.
“FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech”
“Joint Prediction and Denoising for Large-Scale Multilingual Self-Supervised Learning”
William Chen et al · 2023
Later among the works it cites.
“Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data”
Yifan Peng et al · 2023
Later among the works it cites.
“E-branchformer: Branchformer with enhanced merging for speech recognition”
Kwangyoun Kim et al · 2023
Later among the works it cites.
“SLUE-PERB: A Spoken Language Understanding Performance Benchmark and Toolkit”
Siddhant Arora et al · 2023
Later among the works it cites.
“A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks”
Yifan Peng et al · 2023
Later among the works it cites.
“ReazonSpeech: A Free and Massive Corpus for Japanese ASR”, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau · 2022
Cited alongside, same era.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”
Tri Dao et al · 2022
Cited alongside, same era.
“A Study on the Integration of Pre-trained SSL, ASR, LM and SLU Models for Spoken Language Understanding”
Yifan Peng et al · 2022
Cited alongside, same era.
“Emergent Abilities of Large Language Models”
Jason Wei et al · 2022
Cited alongside, same era.
“Distilhubert: Speech Representation Learning by Layer-Wise Distillation of Hidden-Unit Bert”
Heng-Jui Chang, Shu-wen Yang and Hung-yi Lee · 2022
Cited alongside, same era.
“Robust Speech Recognition via Large-Scale Weak Supervision”
Alec Radford et al · 2023
Cited alongside, same era.
“Google usm: Scaling automatic speech recognition beyond 100 languages”
Yu Zhang et al · 2023
Cited alongside, same era.
Yue Yin and Daijiro Mori · 2023
Later among the works it cites.
“Yodas: Youtube-Oriented Dataset for Audio and Speech”
Xinjian Li et al · 2023
Later among the works it cites.
“Structured Pruning of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding”
Yifan Peng et al · 2023
Later among the works it cites.
“DPHuBERT: Joint Distillation and Pruning of Self-Supervised Speech Models”
Yifan Peng, Yui Sudo, Shakeel Muhammad and Shinji Watanabe · 2023
Later among the works it cites.
“I3D: Transformer Architectures with Input-Dependent Dynamic Depth for Speech Recognition”
Yifan Peng, Jaesong Lee and Shinji Watanabe · 2023
Later among the works it cites.
“Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling”
Sanchit Gandhi, Patrick von Platen and Alexander Rush · 2023
Later among the works it cites.
Siddhant Arora et al · 2023
Later among the works it cites.
“SLM: Bridge the Thin Gap Between Speech and Text Foundation Models”
Mingqiu Wang et al · 2023
Later among the works it cites.
“Salmonn: Towards generic hearing abilities for large language models”
Changli Tang et al · 2023
Later among the works it cites.
“Olmo: Accelerating the science of language models”
Dirk Groeneveld et al · 2024
Closest in time.