Fetching the paper…
Reading the bibliography…
Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006 · 2006
Earlier work this paper cites.
Simultaneous translation of lectures and speeches
Christian Fügen, Alex Waibel, and Muntsin Kolss. 2007 · 2007
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Earlier work this paper cites.
Can neural machine translation do simultaneous translation?
Kyunghyun Cho and Masha Esipova. 2016 · 2016
Earlier work this paper cites.
Learning to translate in real-time with neural machine translation
Jiatao Gu, Graham Neubig, Kyunghyun Cho, and Victor O.K. Li. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor O.K. Li, and Richard Socher. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Monotonic Infinite Lookback Attention for Simultaneous Machine Translation
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, and Colin Raffel. 2019 · 2019
Earlier work this paper cites.
Direct Speech-to-Speech Translation with a Sequence-to-Sequence Model
Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Earlier work this paper cites.
STACL: Simultaneous translation with implicit anticipation and controllable latency using prefix-to-prefix framework
Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, and Haifeng Wang. 2019 · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
Efficient Wait-k Models for Simultaneous Machine Translation
Maha Elbayad, Laurent Besacier, and Jakob Verbeek. 2020 · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang. 2020 · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020 · 2020
Earlier work this paper cites.
SIMULEVAL: An evaluation toolkit for simultaneous translation
Xutai Ma, Mohammad Javad Dousti, Changhan Wang, Jiatao Gu, and Juan Pino. 2020a · 2020
Cited alongside, same era.
SimulSpeech: End-to-end simultaneous speech to text translation
Yi Ren, Jinglin Liu, Xu Tan, Chen Zhang, Tao Qin, Zhou Zhao, and Tie-Yan Liu. 2020 · 2020
Cited alongside, same era.
Non-autoregressive machine translation with latent alignments
Chitwan Saharia, William Chan, Saurabh Saxena, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
Covost 2: A massively multilingual speech-to-text translation corpus
Changhan Wang, Anne Wu, and Juan Pino. 2020 · 2020
Cited alongside, same era.
Simultaneous translation policies: From fixed to adaptive
Baigong Zheng, Kaibo Liu, Renjie Zheng, Mingbo Ma, Hairong Liu, and Liang Huang. 2020 · 2020
Cited alongside, same era.
Gaussian multi-head attention for simultaneous machine translation
Shaolei Zhang and Yang Feng. 2022a · 2022
Later among the works it cites.
Information-transport-based policy for simultaneous translation
Shaolei Zhang and Yang Feng. 2022b · 2022
Later among the works it cites.
Wait-info policy: Balancing source and target at information level for simultaneous machine translation
Shaolei Zhang, Shoutao Guo, and Yang Feng. 2022b · 2022
Later among the works it cites.
Liam Dugan, Anshul Wadhawan, Kyle Spence, Chris Callison-Burch, Morgan McGuire, and Victor Zordan. 2023 · 2023
Later among the works it cites.
Daspeech: Directed acyclic transformer for fast and high-quality speech-to-speech translation
Qingkai Fang, Yan Zhou, and Yang Feng. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Direct simultaneous speech-to-text translation assisted by synchronized streaming ASR
Junkun Chen, Mingbo Ma, Renjie Zheng, and Liang Huang. 2021 · 2021
Cited alongside, same era.
A general multi-task learning framework to leverage text data for speech to text tasks
Yun Tang, Juan Pino, Changhan Wang, Xutai Ma, and Dmitriy Genzel. 2021b · 2021
Cited alongside, same era.
RealTranS: End-to-end simultaneous speech translation with convolutional weighted-shrinking transformer
Xingshan Zeng, Liangyou Li, and Qun Liu. 2021 · 2021
Cited alongside, same era.
ICT’s system for AutoSimTrans 2021: Robust char-level simultaneous translation
Shaolei Zhang and Yang Feng. 2021a · 2021
Cited alongside, same era.
Universal simultaneous machine translation with mixture-of-experts wait-k policy
Shaolei Zhang and Yang Feng. 2021b · 2021
Cited alongside, same era.
Future-guided incremental transformer for simultaneous translation
Shaolei Zhang, Yang Feng, and Liangyou Li. 2021 · 2021
Cited alongside, same era.
Learning when to translate for streaming speech
Qian Dong, Yaoming Zhu, Mingxuan Wang, and Lei Li. 2022 · 2022
Cited alongside, same era.
Later among the works it cites.
Adapting offline speech translation models for streaming with future-aware distillation and inference
Biao Fu, Minpeng Liao, Kai Fan, Zhongqiang Huang, Boxing Chen, Yidong Chen, and Xiaodong Shi. 2023 · 2023
Later among the works it cites.
Simultaneous machine translation with tailored reference
Shoutao Guo, Shaolei Zhang, and Yang Feng. 2023b · 2023
Later among the works it cites.
UnitY: Two-pass direct speech-to-speech translation with discrete units
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, and Juan Pino. 2023 · 2023
Later among the works it cites.
Average token delay: A latency metric for simultaneous translation
Yasumasa Kano, Katsuhito Sudoh, and Satoshi Nakamura. 2023 · 2023
Later among the works it cites.
Non-autoregressive streaming transformer for simultaneous translation
Zhengrui Ma, Shaolei Zhang, Shoutao Guo, Chenze Shao, Min Zhang, and Yang Feng. 2023 · 2023
Later among the works it cites.
Attention as a guide for simultaneous speech translation
Sara Papi, Matteo Negri, and Marco Turchi. 2023 · 2023
Later among the works it cites.
Proceedings of the 20th International Conference on Spoken Language Translation (IWSLT 2023) . Association for Computational Linguistics, Toronto, Canada (in-person and online)
Elizabeth Salesky, Marcello Federico, and Marine Carpuat, editors. 2023 · 2023
Later among the works it cites.
Shaolei Zhang, Qingkai Fang, Zhuocheng Zhang, Zhengrui Ma, Yan Zhou, Langlin Huang, Mengyu Bu, Shangtong Gui, Yunji Chen, Xilin Chen, and Yang Feng. 2023 · 2023
Later among the works it cites.
End-to-end simultaneous speech translation with differentiable segmentation
Shaolei Zhang and Yang Feng. 2023a · 2023
Later among the works it cites.
Glancing future for simultaneous machine translation
Shoutao Guo, Shaolei Zhang, and Yang Feng. 2024b · 2024
Closest in time.
A non-autoregressive generation framework for end-to-end simultaneous speech-to-any translation
Zhengrui Ma, Qingkai Fang, Shaolei Zhang, Shoutao Guo, Yang Feng, and Min Zhang. 2024 · 2024
Closest in time.
Hello gpt-4o
OpenAI. 2024 · 2024
Closest in time.