Fetching the paper…
Reading the bibliography…
Impressive image captioning results are achieved in domains with plenty of training image and sentence pairs (e.g., MSCOCO).
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Earlier work this paper cites.
Domain-adversarial neural networks
H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, and M. Marchand · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Y. Kim · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. S. Zemel · 2015
Earlier work this paper cites.
Learning transferable features with deep adaptation networks
M. Long, Y. Cao, J. Wang, and M. I. Jordan · 2015
Cited alongside, same era.
Sequence to sequence-video to text
S. Venugopalan, M. Rohrbach, J. Donahue, R. Mooney, T. Darrell, and K. Saenko · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Cited alongside, same era.
Guided open vocabulary image captioning with constrained beam search
A new dataset and benchmark on animated gif description
Y. Li, Y. Song, L. Cao, J. Tetreault, L. Goldberg, A. Jaimes, and J. Luo · 2016
Later among the works it cites.
Optimization of image description metrics using policy gradient methods
S. Liu, Z. Zhu, N. Ye, S. Guadarrama, and K. Murphy · 2016
Later among the works it cites.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2016
Later among the works it cites.
Learning deep representations of fine-grained visual descriptions
S. Reed, Z. Akata, H. Lee, and B. Schiele · 2016
Later among the works it cites.
Self-critical sequence training for image captioning
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Cited alongside, same era.
Domain-adversarial training of neural networks
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky · 2016
Cited alongside, same era.
Professor forcing: A new algorithm for training recurrent networks
A. Goyal, A. Lamb, Y. Zhang, S. Zhang, A. C. Courville, and Y. Bengio · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Generating visual explanations
L. A. Hendricks, Z. Akata, M. Rohrbach, J. Donahue, B. Schiele, and T. Darrell · 2016
Cited alongside, same era.
Deep compositional captioning: Describing novel object categories without paired training data
L. A. Hendricks, S. Venugopalan, M. Rohrbach, R. Mooney, S. Kate, and T. Darrell · 2016
Cited alongside, same era.
Fcns in the wild: Pixel-level adversarial and constraint-based adaptation
J. Hoffman, D. Wang, F. Yu, and T. Darrell · 2016
Cited alongside, same era.
S. Venugopalan, L. A. Hendricks, M. Rohrbach, R. J. Mooney, T. Darrell, and K. Saenko · 2016
Later among the works it cites.
Title generation for user generated videos
K.-H. Zeng, T.-H. Chen, J. C. Niebles, and M. Sun · 2016
Later among the works it cites.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2017
Closest in time.
No more discrimination: cross city adaptation of road scene segmenters
Y.-H. Chen, W.-Y. Chen, Y.-T. Chen, B.-C. Tsai, Y.-C. F. Wang, and M. Sun · 2017
Closest in time.
Image-to-image translation with conditional adversarial networks
P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros · 2017
Closest in time.
Seqgan: sequence generative adversarial nets with policy gradient
L. Yu, W. Zhang, J. Wang, and Y. Yu · 2017
Closest in time.
Central moment discrepancy (CMD) for domain-invariant representation learning
W. Zellinger, T. Grubinger, E. Lughofer, T. Natschläger, and S. Saminger-Platz · 2017
Closest in time.