Fetching the paper…
Reading the bibliography…
We investigate end-to-end speech-to-text translation on a corpus of audiobooks specifically augmented for this task.
“YAAFE, an Easy to Use and Efficient Audio Feature Extraction Software,”
Benoit Mathieu, Slim Essid, Thomas Fillon, Jacques Prado, and Gaël Richard, · 2010
Earlier work this paper cites.
“Improved Speech-to-Text Translation with the Fisher and Callhome Spanish-English Speech Translation Corpus,”
Matt Post, Gaurav Kumar, Adam Lopez, Damianos Karakos, Chris Callison-Burch, and Sanjeev Khudanpur, · 2013
Earlier work this paper cites.
“Recurrent Neural Network Regularization,”
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals, · 2014
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization,”
Diederik Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Variational dropout and the local reparameterization trick,”
Diederik P Kingma, Tim Salimans, and Max Welling, · 2015
Cited alongside, same era.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems,”
Martin Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Man, Rajat Monga, Sherry Moore, Derek Murray, Jon Shlens, Benoit Steiner, Ilya Sutskever, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Oriol Vinyals, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng, · 2015
Cited alongside, same era.
“An Attentional Model for Speech Translation Without Transcription,”
Long Duong, Antonios Anastasopoulos, David Chiang, Steven Bird, and Trevor Cohn, · 2016
Cited alongside, same era.
“Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation,”
Alexandre Bérard, Olivier Pietquin, Laurent Besacier, and Christophe Servan, · 2016
Cited alongside, same era.
“Breaking the Unwritten Language Barrier: The Bulb Project,”
Gilles Adda, Sebastian Stücker, Martine Adda-Decker, Odette Ambouroue, Laurent Besacier, David Blachon, Hélène Bonneau-Maynard, Pierre Godard, Fatima Hamlaoui, Dmitri Idiatov, Guy-Noël Kouarata, Lori Lamel, Emmanuel-Moselly Makasso, Annie Rialland, Mark Van de Velde, François Yvon, and Sabine Zerbian, · 2016
“Neural Machine Translation of Rare Words with Subword Units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Later among the works it cites.
“Multi-task Sequence to Sequence Learning,”
Minh-Thang Luong, Quoc V Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser, · 2016
Later among the works it cites.
“Sequence-to-Sequence Models Can Directly Transcribe Foreign Speech,”
Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen, · 2017
Later among the works it cites.
“A case study on using speech-to-translation alignments for language documentation,”
Antonios Anastasopoulos and David Chiang, · 2017
Later among the works it cites.
“Nematus: a Toolkit for Neural Machine Translation,”
Rico Sennrich, Orhan Firat, Kyunghyun Cho, Alexandra Birch, Barry Haddow, Julian Hitschler, Marcin Junczys-Dowmunt, Samuel Laeubli, Antonio Valerio, Antonio Valerio Miceli Barone, Jozef Mokry, and Maria Nadejde, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“End-to-End Attention-based Large Vocabulary Speech Recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Cited alongside, same era.
“Listen, Attend and Spell,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation,”
Ali Can Kocabiyikoglu, Laurent Besacier, and Olivier Kraif, · 2018
Closest in time.