Fetching the paper…
Reading the bibliography…
Most of the internet today is composed of digital media that includes videos and images.
Video google: a text retrieval approach to object matching in videos
J. Sivic and A. Zisserman. 2003 · 2003
Earlier work this paper cites.
Distinctive image features from scale-invariant keypoints
David G. Lowe. 2004 · 2004
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, MarcAurelio Ranzato, and Tomas Mikolov. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, and Serge J. Belongie. 2014 · 2014
Cited alongside, same era.
Contextual query expansion for image retrieval
H. Xie, Y. Zhang, J. Tan, L. Guo, and J. Li. 2014 · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Ryan Kiros, Yukun Zhu, and Ruslan Salakhutdinov. 2015 · 2015
Later among the works it cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, and Ryan Kiros. 2015 · 2015
Later among the works it cites.
Recent advance in content-based image retrieval: A literature survey
Wengang Zhou, Houqiang Li, and Qi Tian. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…