Fetching the paper…
Reading the bibliography…
In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos.
Janus: a speech-to-speech translation system using connectionist and symbolic processing strategies
Alex Waibel, Ajay N Jain, Arthur E McNair, Hiroaki Saito, Alexander G Hauptmann, and Joe Tebelskis · 1991
Earlier work this paper cites.
Janus-iii: Speech-to-speech translation in multiple languages
Alon Lavie, Alex Waibel, Lori Levin, Michael Finke, Donna Gates, Marsal Gavalda, Torsten Zeppenfeld, and Puming Zhan · 1997
Earlier work this paper cites.
Ssml: A speech synthesis markup language
Paul Taylor and Amy Isard · 1997
Earlier work this paper cites.
Face translation: A multimodal translation agent
Max Ritter, Uwe Meier, Jie Yang, and Alex Waibel · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn · 2005
Earlier work this paper cites.
Simultaneous translation of lectures and speeches
Christian Fügen, Alex Waibel, and Muntsin Kolss · 2007
Earlier work this paper cites.
Simultaneous german-english lecture translation
Muntsin Kolss, Matthias Wölfel, Florian Kraft, Jan Niehues, Matthias Paulik, and Alex Waibel · 2008
Earlier work this paper cites.
Jibbigo: Speech-to-speech translation on mobile devices
Matthias Eck, Ian Lane, Ying Zhang, and Alex Waibel · 2010
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico · 2012
Earlier work this paper cites.
Simultaneous translation of open domain lectures and speeches, Jan. 3 2012
Alexander Waibel and Christian Fuegen · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Out of time: automated lip sync in the wild
Joon Son Chung and Andrew Zisserman · 2016
Earlier work this paper cites.
Lecture translator-speech translation framework for simultaneous lecture translation
Markus Müller, Thai-Son Nguyen, Jan Niehues, Eunah Cho, Bastian Krüger, Thanh-Le Ha, Kevin Kilgour, Matthias Sperber, Mohammed Mediani, Sebastian Stüker, et al · 2016
Earlier work this paper cites.
Hybrid, offline/online speech translation system, Aug. 30 2016
Naomi Aoki Waibel, Alexander Waibel, Christian Fuegen, and Kay Rottman · 2016
Earlier work this paper cites.
Lip reading sentences in the wild
J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman · 2017
Earlier work this paper cites.
Lip reading in profile
J. S. Chung and A. Zisserman · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
ObamaNet: Photo-realistic lip-sync from text
Rithesh Kumar, Jose Sotelo, Kundan Kumar, Alexandre de Brebisson, and Yoshua Bengio · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger · 2017
Earlier work this paper cites.
Synthesizing Obama: learning lip sync from audio
Supasorn Suwajanakorn, Steven M. Seitz, and Ira Kemelmacher-Shlizerman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Cited alongside, same era.
Deep audio-visual speech recognition
T. Afouras, J. S. Chung, A. Senior, O. Vinyals, and A. Zisserman · 2018
Cited alongside, same era.
Lip movements generation at a glance
Lele Chen, Zhiheng Li, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2018
Cited alongside, same era.
Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Estève · 2018
Cited alongside, same era.
Low-latency neural speech translation
Jan Niehues, Ngoc-Quan Pham, Thanh-Le Ha, Matthias Sperber, and Alex Waibel · 2018
Cited alongside, same era.
Super-human performance in online low-latency recognition of conversational speech
Thai-Son Nguyen, Sebastian Stüker, and Alex Waibel · 2020
Later among the works it cites.
Relative Positional Encoding for Speech Recognition and Direct Translation
Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen, Thai-Son Nguyen, Elizabeth Salesky, Sebastian St üker, Jan Niehues, and Alex Waibel · 2020
Later among the works it cites.
A lip sync expert is all you need for speech to lip generation in the wild
KR Prajwal, Rudrabha Mukhopadhyay, Vinay P Namboodiri, and CV Jawahar · 2020
Later among the works it cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2020
Later among the works it cites.
Neural Voice Puppetry: Audio-Driven Facial Reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, and Matthias Nießner · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How2: a large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
X2face: A network for controlling face generation using images, audio, and pose codes
Olivia Wiles, A Koepke, and Andrew Zisserman · 2018
Cited alongside, same era.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Cited alongside, same era.
Hierarchical cross-modal talking face generation with dynamic pixel-wise loss
Lele Chen, Ross K Maddox, Zhiyao Duan, and Chenliang Xu · 2019
Cited alongside, same era.
Ju-chieh Chou, Cheng-chieh Yeh, and Hung-yi Lee · 2019
Cited alongside, same era.
MuST-C: a Multilingual Speech Translation Corpus
Mattia A. Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi · 2019
Cited alongside, same era.
Realistic Speech-Driven Facial Animation with GANs
Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic · 2020
Later among the works it cites.
MakeltTalk: speaker-aware talking-head animation
Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li · 2020
Later among the works it cites.
Must-c: A multilingual corpus for end-to-end speech translation
Roldano Cattoni, Mattia Antonino Di Gangi, Luisa Bentivogli, Matteo Negri, and Marco Turchi · 2021
Later among the works it cites.
Ad-nerf: Audio driven neural radiance fields for talking head synthesis
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang · 2021
Later among the works it cites.
Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization
Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, and Chris Bregler · 2021
Later among the works it cites.
Isometricmt: Neural machine translation for automatic dubbing
Surafel M Lakew, Yogesh Virkar, Prashant Mathur, and Marcello Federico · 2021
Later among the works it cites.
Fragmentvc: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention, 2021
Yist Y. Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung yi Lee, and Lin shan Lee · 2021
Later among the works it cites.
KIT’s IWSLT 2021 offline speech translation system
Tuan Nam Nguyen, Thai Son Nguyen, Christian Huber, Ngoc-Quan Pham, Thanh-Le Ha, Felix Schneider, and Sebastian Stüker · 2021
Later among the works it cites.
Multilingual speech translation KIT @ IWSLT2021
Ngoc-Quan Pham, Tuan Nam Nguyen, Thanh-Le Ha, Sebastian Stüker, Alexander Waibel, and Dan He · 2021
Later among the works it cites.
Efficient Weight Factorization for Multilingual Speech Recognition
Ngoc-Quan Pham, Tuan-Nam Nguyen, Sebastian Stüker, and Alex Waibel · 2021
Later among the works it cites.
Vqmivc: Vector quantization and mutual information-based unsupervised speech representation disentanglement for one-shot voice conversion, 2021
Disong Wang, Liqun Deng, Yu Ting Yeung, Xiao Chen, Xunying Liu, and Helen Meng · 2021
Later among the works it cites.
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu · 2021
Later among the works it cites.
Facial: Synthesizing dynamic talking face with implicit attribute learning
Chenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng, Saifeng Ni, Madhukar Budagavi, and Xiaohu Guo · 2021
Later among the works it cites.
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan · 2021
Later among the works it cites.
Pose-controllable talking face generation by implicitly modularized audio-visual representation
Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu · 2021
Later among the works it cites.
Findings of the iwslt 2022 evaluation campaign
Antonios Anastasopoulos, Loïc Barrault, Luisa Bentivogli, Marcely Zanon Boito, Ondřej Bojar, Roldano Cattoni, Anna Currey, Georgiana Dinu, Kevin Duh, Maha Elbayad, et al · 2022
Closest in time.
Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation
Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou · 2022
Closest in time.
Adaptive multilingual speech recognition with pretrained models
Ngoc-Quan Pham, Alex Waibel, and Jan Niehues · 2022
Closest in time.
DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering
Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai, and Xiaokang Yang · 2022
Closest in time.