Fetching the paper…
Reading the bibliography…
In this paper, we address the task of learning novel visual concepts, and their interactions with other concepts, from a few images with sentence descriptions.
Acquiring a single new word
S. Carey and E. Bartlett · 1978
Earlier work this paper cites.
Word learning in children: An examination of fast mapping
T. H. Heibeck and E. M. Markman · 1987
Earlier work this paper cites.
Finding structure in time
J. L. Elman · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
How children learn the meanings of words
P. Bloom · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
One-shot learning of object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2006
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgements
A. Lavie and A. Agarwal · 2007
Earlier work this paper cites.
Learning deep architectures for ai
Y. Bengio · 2009
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
One-shot learning with a hierarchical nonparametric bayesian model
R. Salakhutdinov, J. Tenenbaum, and A. Torralba · 2010
Earlier work this paper cites.
Fast mapping and slow mapping in children’s word learning
D. Swingley · 2010
Earlier work this paper cites.
Large scale image annotation: learning to rank with joint word-image embeddings
J. Weston, S. Bengio, and N. Usunier · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
One shot learning of simple visual concepts
B. M. Lake, R. Salakhutdinov, J. Gross, and J. B. Tenenbaum · 2011
Earlier work this paper cites.
From image annotation to image description
A. Gupta and P. Mannem · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Earlier work this paper cites.
Midge: Generating image descriptions from computer vision detections
M. Mitchell, X. Han, J. Dodge, A. Mensch, A. Goyal, A. Berg, K. Yamaguchi, T. Berg, K. Stratos, and H. Daumé III · 2012
Earlier work this paper cites.
Augmented attribute representations
V. Sharmanska, N. Quadrianto, and C. H. Lampert · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2013
Cited alongside, same era.
Write a classifier: Zero-shot learning using purely textual descriptions
M. Elhoseiny, B. Saleh, and A. Elgammal · 2013
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Cited alongside, same era.
Recurrent continuous translation models
N. Kalchbrenner and P. Blunsom · 2013
Cited alongside, same era.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Later among the works it cites.
A pooling approach to modelling spatial relations for image retrieval and annotation
M. Malinowski and M. Fritz · 2014
Later among the works it cites.
Explain images with multimodal recurrent neural networks
J. Mao, W. Xu, Y. Yang, J. Wang, and A. L. Yuille · 2014
Later among the works it cites.
Deepid-net: multi-stage and deformable deep convolutional neural networks for object detection
W. Ouyang, P. Luo, X. Zeng, S. Qiu, Y. Tian, H. Li, S. Yang, Z. Wang, Y. Xiong, C. Qian, et al · 2014
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge, 2014
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Zero-shot learning through cross-modal transfer
R. Socher, M. Ganjoo, C. D. Manning, and A. Ng · 2013
Cited alongside, same era.
Zero-shot learning via visual abstraction
S. Antol, C. L. Zitnick, and D. Parikh · 2014
Cited alongside, same era.
Learning a recurrent visual representation for image caption generation
X. Chen and C. L. Zitnick · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2014
Cited alongside, same era.
Comparing automatic evaluation measures for image description
D. Elliott and F. Keller · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Later among the works it cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, Q. Le, C. Manning, and A. Ng · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Learning categories from few examples with multi model knowledge transfer
T. Tommasi, F. Orabona, and B. Caputo · 2014
Later among the works it cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2014
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Later among the works it cites.
Learning from weakly supervised data by the expectation loss svm (e-svm) algorithm
J. Zhu, J. Mao, and A. L. Yuille · 2014
Later among the works it cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Closest in time.
Microsoft coco captions: Data collection and evaluation server
X. Chen, H. Fang, T. Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick · 2015
Closest in time.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Closest in time.
Are you talking to a machine? dataset and methods for multilingual image question answering
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Closest in time.
Multimodal convolutional neural networks for matching image and sentence
L. Ma, Z. Lu, L. Shang, and H. Li · 2015
Closest in time.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, Z. Huang, and A. Yuille · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, C. A. Cho, Kyunghyun, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Closest in time.