Fetching the paper…
Reading the bibliography…
Automatic image captioning has improved significantly over the last few years, but the problem is far from being solved, with state of the art models still often producing low quality captions when used in the wild.
Graph-rise: Graph-regularized image semantic embedding
Da-Cheng Juan, Chun-Ta Lu, Zhen Li, Futang Peng, Aleksei Timofeev, Yi-Ting Chen, Yaxi Gao, Tom Duerig, Andrew Tomkins, and Sujith Ravi. 2019 · 1902
Earlier work this paper cites.
Estimating the sentence-level quality of machine translation systems
Lucia Specia, Nicola Cancedda, Marc Dymetman, Marco Turchi, and Nello Cristianini. 2009 · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Trustrank: Inducing trust in automatic translations via ranking
Radu Soricut and Abdessamad Echihabi. 2010 · 2010
Earlier work this paper cites.
Crowd-sourcing of human judgments of machine translation fluency
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
QuEst - A Translation Quality Estimation Framework
Lucia Specia, Kashif Shah, Jose Guilherme Camargo de Souza, and Trevor Cohn. 2013 · 2013
Earlier work this paper cites.
Adequacy–fluency metrics: Evaluating mt in the continuous space model framework
R. E. Banchs, L. F. D’Haro, and H. Li. 2015 · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Quality estimation from scratch (quetch): Deep learning for word-level translation quality estimation
Julia Kreutzer, Shigehiko Schamoni, and Stefan Riezler. 2015 · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2016 · 2015
Cited alongside, same era.
Empirical evaluation of rectified activations in convolutional network
Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. 2015 · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
A recurrent neural networks approach for estimating the quality of machine translation output
Hyun Kim and Jong-Hyeok Lee. 2016 · 2016
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Universal sentence encoder for English
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018 · 2018
Later among the works it cites.
Learning to evaluate image captioning
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge J. Belongie. 2018 · 2018
Later among the works it cites.
Illustrative language understanding: Large-scale visual grounding with image search
Jamie Kiros, William Chan, and Geoffrey Hinton. 2018 · 2018
Later among the works it cites.
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari. 2018 · 2018
Later among the works it cites.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christian Szegedy, Sergey Ioffe, and Vincent Vanhoucke. 2016 · 2016
Cited alongside, same era.
Predictor-estimator: Neural quality estimation based on target word prediction for machine translation
Hyun Kim, Hun-Young Jung, Hongseok Kwon, Jong-Hyeok Lee, and Seung-Hoon Na. 2017 · 2017
Cited alongside, same era.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael Bernstein, and Li Fei-Fei. 2017 · 2017
Cited alongside, same era.
Understanding blind people’s experiences with computer-generated captions of social media images
Haley MacLeod, Cynthia L. Bennett, Meredith Ringel Morris, and Edward Cutrell. 2017 · 2017
Cited alongside, same era.
Pushing the limits of translation quality estimation
André F. T. Martins, Marcin Junczys-Dowmunt, Fábio Kepler, Ramón Fernández Astudillo, Chris Hokamp, and Roman Grundkiewicz. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Findings of the wmt 2018 shared task on quality estimation
Lucia Specia, Frederic Blain, Varvara Logacheva, Ramon Astudillo, and Andre Martins. 2019 · 2018
Later among the works it cites.
Alibaba submission for wmt18 quality estimation task
Jiayi Wang, Kai Fan, Bo Li, Fengming Zhou, Boxing Chen, Yangbin Shi, and Luo Si. 2018 · 2018
Later among the works it cites.
Decoupled box proposal and featurization with ultrafine-grained semantic labels improve image captioning and visual question answering
Soravit Changpinyo, Bo Pang, Piyush Sharma, and Radu Soricut. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Cross-modal coherence modeling for caption generation
Malihe Alikhani, Piyush Sharma, Shengjie Li, Radu Soricut, and Matthew Stone. 2020 · 2020
Closest in time.
Reinforcing an image caption generator using off-line human feedback
Paul Hongsuck Seo, Piyush Sharma, Tomer Levinboim, and Radu Soricut. 2020 · 2020
Closest in time.