Fetching the paper…
Reading the bibliography…
Image captioning has focused on generalizing to images drawn from the same distribution as the training set, and not to the more challenging problem of generalizing to different distributions of images.
Has a consensus NL generation architecture appeared, and is it psycholinguistically plausible?
Ehud Reiter. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Building applied natural language generation systems
Ehud Reiter and Robert Dale. 1997 · 1997
Earlier work this paper cites.
What the eyes say about speaking
Zenzi M Griffin and Kathryn Bock. 2000 · 2000
Earlier work this paper cites.
Introduction to the CoNLL-2000 shared task chunking
Erik F. Tjong Kim Sang and Sabine Buchholz. 2000 · 2000
Earlier work this paper cites.
The compositionality papers
Jerry A. Fodor and Ernie Lepore. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Combinatory categorial grammar
M. Steedman and J. Baldridge. 2006 · 2006
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Scan patterns predict sentence production in the cross-modal processing of visual scenes
Moreno I Coco and Frank Keller. 2012 · 2012
Earlier work this paper cites.
Collective generation of natural image descriptions
Polina Kuznetsova, Vicente Ordonez, Alexander Berg, Tamara Berg, and Yejin Choi. 2012 · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
Margaret Mitchell, Jesse Dodge, Amit Goyal, Kota Yamaguchi, Karl Stratos, Xufeng Han, Alyssa Mensch, Alex Berg, Tamara Berg, and Hal Daumé III. 2012 · 2012
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Treetalk: Composition and compression of trees for image descriptions
Polina Kuznetsova, Vicente Ordonez, Tamara L. Berg, and Yejin Choi. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané. 2016 · 2016
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Learning to generalize to new compositions in image understanding
Yuval Atzmon, Jonathan Berant, Vahid Kezami, Amir Globerson, and Gal Chechik. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Neural baby talk
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018 · 2018
Later among the works it cites.
Measuring the diversity of automatic image descriptions
Emiel van Miltenburg, Desmond Elliott, and Piek Vossen. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Universal dependency parsing from scratch
Peng Qi, Timothy Dozat, Yuhao Zhang, and Christopher D. Manning. 2018 · 2018
Later among the works it cites.
A multi-task learning approach for image captioning
Wei Zhao, Benyou Wang, Jianbo Ye, Min Yang, Zhou Zhao, Ruotian Luo, and Yu Qiao. 2018 · 2018
Later among the works it cites.
Linguistic generalization and compositionality in modern artificial neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017 · 2017
Cited alongside, same era.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy. 2017 · 2017
Cited alongside, same era.
From red wine to red tomato: Composition with context
Ishan Misra, Abhinav Gupta, and Martial Hebert. 2017 · 2017
Cited alongside, same era.
Predicting target language CCG supertags improves neural machine translation
Maria Nădejde, Siva Reddy, Rico Sennrich, Tomasz Dwojak, Marcin Junczys-Dowmunt, Philipp Koehn, and Alexandra Birch. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Skeleton key: Image captioning by skeleton-attribute decomposition
Yufei Wang, Zhe Lin, Xiaohui Shen, Scott Cohen, and Garrison W Cottrell. 2017 · 2017
Cited alongside, same era.
A* CCG parsing with a supertag and dependency factored model
Masashi Yoshikawa, Hiroshi Noji, and Yuji Matsumoto. 2017 · 2017
Cited alongside, same era.
Marco Baroni. 2019 · 2019
Later among the works it cites.
Incorporating source syntax into transformer-based neural machine translation
Anna Currey and Kenneth Heafield. 2019 · 2019
Later among the works it cites.
Fast, diverse and accurate image captioning guided by part-of-speech
Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander G Schwing, and David Forsyth. 2019 · 2019
Later among the works it cites.
A comprehensive survey of deep learning for image captioning
MD. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga. 2019 · 2019
Later among the works it cites.
Joint syntax representation learning and visual cue translation for video captioning
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia. 2019 · 2019
Later among the works it cites.
Meta learning for image captioning
Nannan Li, Zhenzhong Chen, and Shan Liu. 2019 · 2019
Later among the works it cites.
Compositional generalization in image captioning
Mitja Nikolaus, Mostafa Abdou, Matthew Lamm, Rahul Aralikatte, and Desmond Elliott. 2019 · 2019
Later among the works it cites.
Meshed-memory transformer for image captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 2020
Later among the works it cites.
Normalized and geometry-aware self-attention network for image captioning
Longteng Guo, Jing Liu, Xinxin Zhu, Peng Yao, Shichen Lu, and Hanqing Lu. 2020 · 2020
Later among the works it cites.
Att-bm-som: A framework of effectively choosing image information and optimizing syntax for image captioning
Zhenyu Yang and Qiao Liu. 2020 · 2020
Later among the works it cites.
Improving image captioning evaluation by considering inter references variance
Yanzhi Yi, Hangyu Deng, and Jinglu Hu. 2020 · 2020
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.
Deep recursive neural networks for compositionality in language
Ozan Irsoy and Claire Cardie. 2014 · 2096
Closest in time.