Fetching the paper…
Reading the bibliography…
The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation.
A statistical model-based voice activity detection
Jongseo Sohn, Nam Soo Kim, and Wonyong Sung. 1999 · 1999
Earlier work this paper cites.
BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic Sentence Segmentation and Punctuation Prediction for Spoken Language Translation
Evgeny Matusov, Arne Mauser, and Hermann Ney. 2006 · 2006
Earlier work this paper cites.
A Study of Translation Edit Rate with Targeted Human Annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
LIUM SpkDiarization: An Open Source Toolkit For Diarization
Sylvain Meignier and Teva Merlin. 2010 · 2010
Earlier work this paper cites.
Determining the placement of German verbs in English–to–German SMT
Anita Gojun and Alexander Fraser. 2012 · 2012
Earlier work this paper cites.
A semi-Markov model for speech segmentation with an utterance-break prior
Mark Sinclair, Peter Bell, Alexandra Birch, and Fergus Mcinnes. 2014 · 2014
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
OpenSubtitles2016: Extracting Large Parallel Corpora from Movie and TV Subtitles
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
OPUS – Parallel Corpora for Everyone
Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
MMT: New Open Source MT for the Translation Industry
Nicola Bertoldi, Roldano Cattoni, Mauro Cettolo, Amin Farajian, Marcello Federico, Davide Caroselli, Luca Mastrostefano, Andrea Rossi, Marco Trombetti, Ulrich Germann, and David Madl. 2017 · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation
End-to-End Speech Translation with Knowledge Distillation
Yuchen Liu, Hao Xiong, Jiajun Zhang, Zhongjun He, Hua Wu, Haifeng Wang, and Chengqing Zong. 2019 · 2019
Later among the works it cites.
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
FINDINGS OF THE IWSLT 2020 EVALUATION CAMPAIGN
Ebrahim Ansari, Amittai Axelrod, Nguyen Bach, Ondřej Bojar, Roldano Cattoni, Fahim Dalvi, Nadir Durrani, Marcello Federico, Christian Federmann, Jiatao Gu, Fei Huang, Kevin Knight, Xutai Ma, Ajay Nagesh, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Xing Shi, Sebastian Stüker, Marco Turchi, Alexander Waibel, and Changhan Wang. 2020 · 2020
Later among the works it cites.
Start-Before-End and End-to-End: Neural Speech Translation by AppTek and RWTH Aachen University
Parnia Bahar, Patrick Wilken, Tamer Alkhouli, Andreas Guta, Pavel Golik, Evgeny Matusov, and Christian Herold. 2020 · 2020
Later among the works it cites.
On target segmentation for direct speech translation
Mattia A. Di Gangi, Marco Gaido, Matteo Negri, and Marco Turchi. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Estève. 2018 · 2018
Cited alongside, same era.
Hallucinations in Neural Machine Translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018 · 2018
Cited alongside, same era.
How2: A Large-scale Dataset For Multimodal Language Understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze. 2018 · 2018
Cited alongside, same era.
Findings of the 2019 conference on machine translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
Cited alongside, same era.
Adapting Transformer to End-to-End Spoken Language Translation
Mattia A. Di Gangi, Matteo Negri, and Marco Turchi. 2019 · 2019
Cited alongside, same era.
Leveraging Weakly Supervised Data to Improve End-to-End Speech-to-Text Translation
Ye Jia, Melvin Johnson, Wolfgang Macherey, Ron J. Weiss, Yuan Cao, Chung-Cheng Chiu, Naveen Ari, Stella Laurenzo, and Yonghui Wu. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Europarl-ST: A Multilingual Corpus For Speech Translation Of Parliamentary Debates
Javier Iranzo-Sánchez, Joan A. Silvestre-Cerdà, Javier Jorge, Nahuel Roselló, Giménez. Adrià, Albert Sanchis, Jorge Civera, and Alfons Juan. 2020 · 2020
Later among the works it cites.
Is 42 the Answer to Everything in Subtitling-oriented Speech Translation?
Alina Karakanta, Matteo Negri, and Marco Turchi. 2020 · 2020
Later among the works it cites.
Improving Sequence-to-sequence Speech Recognition Training with On-the-fly Data Augmentation
Thai-Son Nguyen, Sebastian Stueker, Jan Niehues, and Alex Waibel. 2020 · 2020
Later among the works it cites.
SRPOL’s system for the IWSLT 2020 end-to-end speech translation task
Tomasz Potapczyk and Pawel Przybysz. 2020 · 2020
Later among the works it cites.
MuST-C: A multilingual corpus for end-to-end speech translation
Roldano Cattoni, Mattia A. Di Gangi, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021 · 2021
Closest in time.
On Knowledge Distillation for Direct Speech Translation
Marco Gaido, Mattia A. Di Gangi, Matteo Negri, and Marco Turchi. 2020 · 2021
Closest in time.