Fetching the paper…
Reading the bibliography…
While most machine translation systems to date are trained on large parallel corpora, humans learn language in a different way: by being grounded in an environment and interacting with other humans.
Doubly-attentive decoder for multi-modal neural machine translation
Iacer Calixto, Qun Liu, and Nick Campbell · 1924
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Learning bilingual lexicons using the visual similarity of labeled web images
Shane Bergsma and Benjamin Van Durme · 2011
Earlier work this paper cites.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff A. Bilmes, and Karen Livescu · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, and Andrew Y. Ng · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
Learning language through pictures
Grzegorz Chrupala, Ákos Kádár, and Afra Alishahi · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Fei-Fei Li · 2015
Earlier work this paper cites.
Visual bilingual lexicon induction with transferred convnet features
Douwe Kiela, Ivan Vulić, and Stephen Clark · 2015
Earlier work this paper cites.
Multimodal convolutional neural networks for matching image and sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li · 2015
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2015
Cited alongside, same era.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2015
Cited alongside, same era.
Multimodal attention for neural machine translation
Ozan Caglayan, Loïc Barrault, and Fethi Bougares · 2016
Cited alongside, same era.
Correlational neural networks
Sarath Chandar, Mitesh M. Khapra, Hugo Larochelle, and Balaraman Ravindran · 2016
Cited alongside, same era.
Findings of the second shared task on multimodal machine translation and multilingual image description
Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia · 2017
Closest in time.
Emergent language in a multi-modal, multi-step referential game
Katrina Evtimova, Andrew Drozdov, Douwe Kiela, and Kyunghyun Cho · 2017
Closest in time.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Closest in time.
Image Pivoting for Learning Multilingual Multimodal Representations
Spandana Gella, Rico Sennrich, Frank Keller, and Mirella Lapata · 2017
Closest in time.
Emergence of language with multi-agent games: Learning to communicate with sequences of symbols
Serhii Havrylov and Ivan Titov · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia · 2016
Cited alongside, same era.
Zero-resource translation with multi-lingual neural machine translation
Orhan Firat, Baskaran Sankaran, Yaser Al-Onaizan, Fatos T. Yarman-Vural, and Kyunghyun Cho · 2016
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Multimodal pivots for image caption translation
Julian Hitschler, Shigehiko Schamoni, and Stefan Riezler · 2016
Cited alongside, same era.
Learning to play guess who? and inventing a grounded language as a consequence
Emilio Jorge, Mikael Kågebäck, and Emil Gustavsson · 2016
Cited alongside, same era.
Learning visual features from large weakly supervised data
Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache · 2016
Cited alongside, same era.
Bridge correlational neural networks for multilingual multimodal representation learning
Janarthanan Rajendran, Mitesh M. Khapra, Sarath Chandar, and Balaraman Ravindran · 2016
Cited alongside, same era.
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Closest in time.
Learning visually grounded sentence representations
Douwe Kiela, Alexis Conneau, Allan Jabri, and Maximilian Nickel · 2017
Closest in time.
Natural language does not emerge ’naturally’ in multi-agent dialog
Satwik Kottur, José M. F. Moura, Stefan Lee, and Dhruv Batra · 2017
Closest in time.
Multi-agent cooperation and the emergence of (natural) language
Angeliki Lazaridou, Alexander Peysakhovich, and Marco Baroni · 2017
Closest in time.
Deal or no deal? end-to-end learning for negotiation dialogues
Mike Lewis, Denis Yarats, Yann N. Dauphin, Devi Parikh, and Dhruv Batra · 2017
Closest in time.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Closest in time.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J. Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Closest in time.
Emergence of grounded compositional language in multi-agent populations
Igor Mordatch and Pieter Abbeel · 2017
Closest in time.
Zero-resource machine translation by multimodal encoder–decoder network with multimedia pivot
Hideki Nakayama and Noriki Nishida · 2017
Closest in time.
The university of edinburgh’s neural MT systems for WMT17
Rico Sennrich, Alexandra Birch, Anna Currey, Ulrich Germann, Barry Haddow, Kenneth Heafield, Antonio Valerio Miceli Barone, and Philip Williams · 2017
Closest in time.
STAIR captions: Constructing a large-scale japanese image caption dataset
Yuya Yoshikawa, Yutaro Shigeto, and Akikazu Takeuchi · 2017
Closest in time.