Fetching the paper…
Reading the bibliography…
In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations.
Switchboard: Telephone speech corpus for research and development
John J Godfrey, Edward C Holliman, and Jane McDaniel · 1992
Earlier work this paper cites.
Audio visual speech recognition
Chalapathy Neti, Gerasimos Potamianos, Juergen Luettin, Iain Matthews, Herve Glotin, Dimitra Vergyri, June Sison, and Azad Mashari · 2000
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Grounding conceptual knowledge in modality-specific systems
Lawrence W Barsalou, W Kyle Simmons, Aron K Barbey, and Christine D Wilson · 2003
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och · 2004
Earlier work this paper cites.
An audio-visual corpus for speech perception and automatic speech recognition
Martin Cooke, Jon Barker, Stuart Cunningham, and Xu Shao · 2006
Earlier work this paper cites.
The IAPR TC-12 benchmark: A new evaluation resource for visual information systems
Michael Grubinger, Paul D. Clough, Henning Muller, and Thomas Desealers · 2006
Earlier work this paper cites.
Duc in context
Paul Over, Hoa Dang, and Donna Harman · 2007
Earlier work this paper cites.
Building a persistent workforce on Mechanical Turk for multilingual data collection
David L. Chen and William B. Dolan · 2011
Earlier work this paper cites.
Multimodal summarization of complex sentences
Naushad UzZaman, Jeffrey P Bigham, and James F Allen · 2011
Earlier work this paper cites.
The Kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Annotated gigaword
Courtney Napoles, Matthew Gormley, and Benjamin Van Durme · 2012
Earlier work this paper cites.
Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Earlier work this paper cites.
Improved Speech-to-Text translation with the Fisher and Callhome Spanish-English speech translation corpus
Matt Post, Gaurav Kumar, Adam Lopez, Damianos Karakos, Chris Callison-Burch, and Sanjeev Khudanpur · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Shoou-I Yu, Lu Jiang, and Alexander Hauptmann · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micha Hodosh, and Julia Hockenmaier · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Microsoft COCO captions: Data collection and evaluation server
Visual features for context-aware speech recognition
Abhinav Gupta, Yajie Miao, Leonardo Neves, and Florian Metze · 2017
Later among the works it cites.
NMTPY: A flexible toolkit for advanced neural machine translation systems
Ozan Caglayan, Mercedes García-Martínez, Adrien Bardet, Walid Aransa, Fethi Bougares, and Loïc Barrault · 2017
Later among the works it cites.
Sequence-to-sequence models can directly translate foreign speech
Ron J Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen · 2017
Later among the works it cites.
Attention strategies for multi-source sequence-to-sequence learning
Jindřich Libovický and Jindřich Helcl · 2017
Later among the works it cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick · 2015
Cited alongside, same era.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom · 2015
Cited alongside, same era.
A shared task on multimodal machine translation and crosslingual image description (WMT)
Lucia Specia, Stella Frank, Khalil Sima’an, and Desmond Elliott · 2016
Cited alongside, same era.
Open-domain audio-visual speech recognition: A deep learning approach
Yajie Miao and Florian Metze · 2016
Cited alongside, same era.
Tasviret: Görüntülerden otomatik türkçe açıklama oluşturma İçin bir denektaçı veri kümesi (TasvirEt: A benchmark dataset for automatic Turkish description generation from images)
Mesut Erhan Unal, Begum Citamak, Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler Cinbis, and Ruket Cakici · 2016
Cited alongside, same era.
Adding Chinese captions to images
Xirong Li, Weiyu Lan, Jianfeng Dong, and Hailong Liu · 2016
Cited alongside, same era.
Multi30k: Multilingual english-german image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia · 2016
Cited alongside, same era.
Yuya Yoshikawa, Yutaro Shigeto, and Akikazu Takeuchi · 2017
Later among the works it cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Chris Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Later among the works it cites.
Multi-modal summarization for asynchronous collection of text, image, audio and video
Haoran Li, Junnan Zhu, Cong Ma, Jiajun Zhang, and Chengqing Zong · 2017
Later among the works it cites.
Nematus: a toolkit for neural machine translation
Rico Sennrich, Orhan Firat, Kyunghyun Cho, Alexandra Birch-Mayne, Barry Haddow, Julian Hitschler, Marcin Junczys-Dowmunt, Samuel Läubli, Antonio Miceli Barone, Jozef Mokry, and Maria Nadejde · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Later among the works it cites.
End-to-end multimodal speech recognition
Shruti Palaskar, Ramon Sanabria, and Florian Metze · 2018
Closest in time.
End-to-end automatic speech translation of audiobooks
Alexandre Bérard, Laurent Besacier, Ali Can Kocabiyikoglu, and Olivier Pietquin · 2018
Closest in time.
Multimodal abstractive summarization of open-domain videos
Jindřich Libovický, Shruti Palaskar, Spandana Gella, and Florian Metze · 2018
Closest in time.
Findings of the second shared task on multimodal machine translation and multilingual image description
Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia · 2018
Closest in time.
Findings of the shared task on multimodal machine translation (WMT)
Loïc Barrault, Fethi Bougares, Lucia Specia, Chiraag Lala, Desmond Elliott, and Stella. Frank · 2018
Closest in time.
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
Ali Can Kocabiyikoglu, Laurent Besacier, and Olivier Kraif · 2018
Closest in time.
Can spatiotemporal 3d CNNs retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Closest in time.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo · 2018
Closest in time.