Fetching the paper…
Reading the bibliography…
In this paper, we propose Text2Scene, a model that generates various forms of compositional scene representations from natural language descriptions.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Introduction to the conll-2005 shared task: Semantic role labeling
Xavier Carreras and Lluís Màrquez · 2005
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal · 2007
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. Lawrence Zitnick and Devi Parikh · 2013
Earlier work this paper cites.
Learning the visual interpretation of sentences
C. Lawrence Zitnick, Devi Parikh, and Lucy Vanderwende · 2013
Earlier work this paper cites.
Predicting object dynamics in scenes
David F Fouhey and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Treetalk: Composition and compression of trees for image descriptions
Polina Kuznetsova, Vicente Ordonez, Tamara Berg, and Yejin Choi · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Nonparametric method for data-driven image captioning
Rebecca Mason and Eugene Charniak · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
James Philbin Florian Schroff, Dmitry Kalenichenko · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Cited alongside, same era.
Guided open vocabulary image captioning with constrained beam search
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2017
Later among the works it cites.
Photographic image synthesis with cascaded refinement networks
Qifeng Chen and Vladlen Koltun · 2017
Later among the works it cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Later among the works it cites.
Codraw: Visual dialog for collaborative drawing
Jin-Hwa Kim, Devi Parikh, Dhruv Batra, Byoung-Tak Zhang, and Yuandong Tian · 2017
Later among the works it cites.
Commonly uncommon: Semantic sparsity in situation recognition
Mark Yatskar, Vicente Ordonez, Luke Zettlemoyer, and Ali Farhadi · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Convolutional recurrent neural networks: Learning spatial dependencies for image representation
Z. Zuo, B. Shuai, G. Wang, X. Liu, X. Wang, B. Wang, and Y. Chen · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Large scale retrieval and generation of image descriptions
Vicente Ordonez, Xufeng Han, Polina Kuznetsova, Girish Kulkarni, Margaret Mitchell, Kota Yamaguchi, Karl Stratos, Amit Goyal, Jesse Dodge, Alyssa Mensch, et al · 2016
Cited alongside, same era.
Learning what and where to draw
Scott Reed, Zeynep Akata, Santosh Mohan, Samuel Tenka, Bernt Schiele, and Honglak Lee · 2016
Cited alongside, same era.
Xuwang Yin and Vicente Ordonez · 2017
Later among the works it cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaolei Huang, Xiaogang Wang, and Dimitris Metaxas · 2017
Later among the works it cites.
Deep learning for semantic composition
Xiaodan Zhu and Edward Grefenstette · 2017
Later among the works it cites.
Imagine this! scripts to compositions to videos
Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi · 2018
Closest in time.
Coco-stuff: Thing and stuff classes in context
Jasper Uijlings Holger Caesar and Vittorio Ferrari · 2018
Closest in time.
Inferring semantic layout for hierarchical text-to-image synthesis
Seunghoon Hong, Dingdong Yang, Jongwook Choi, and Honglak Lee · 2018
Closest in time.
Image generation from scene graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei · 2018
Closest in time.
Semi-parametric image synthesis
Xiaojuan Qi, Qifeng Chen, Jiaya Jia, and Vladlen Koltun · 2018
Closest in time.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He · 2018
Closest in time.
Photographic text-to-image synthesis with a hierarchically-nested adversarial network
Zizhao Zhang, Yuanpu Xie, and Lin Yang · 2018
Closest in time.