Fetching the paper…
Reading the bibliography…
This paper proposes a new task, MemexQA: given a collection of photos or videos from a user, the goal is to automatically answer questions that help users recover their memory about events captured in the collection.
The atlantic monthly
V. Bush · 1945
Earlier work this paper cites.
Mylifebits: fulfilling the memex vision
J. Gemmell, G. Bell, R. Lueder, S. Drucker, and C. Wong · 2002
Earlier work this paper cites.
Video google: A text retrieval approach to object matching in videos
J. Sivic, A. Zisserman, et al · 2003
Earlier work this paper cites.
A study of smoothing methods for language models applied to information retrieval
C. Zhai and J. Lafferty · 2004
Earlier work this paper cites.
Building watson: An overview of the deepqa project
D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, et al · 2010
Earlier work this paper cites.
Still building the memex
S. Davies · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov and J. Dean · 2013
Earlier work this paper cites.
J. Weston, S. Chopra, and A. Bordes · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Bridging the ultimate semantic gap: A semantic search engine for internet videos
L. Jiang, S.-I. Yu, D. Meng, T. Mitamura, and A. G. Hauptmann · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
M. Malinowski, M. Rohrbach, and M. Fritz · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. C. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
S. Chandar, S. Ahn, H. Larochelle, P. Vincent, G. Tesauro, and Y. Bengio · 2016
Cited alongside, same era.
Gated-attention readers for text comprehension
B. Dhingra, H. Liu, W. W. Cohen, and R. Salakhutdinov · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Reasonet: Learning to stop reading in machine comprehension
Y. Shen, P.-S. Huang, J. Gao, and W. Chen · 2016
Later among the works it cites.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Later among the works it cites.
FVQA: fact-based visual question answering
P. Wang, Q. Wu, C. Shen, A. van den Hengel, and A. R. Dick · 2016
Later among the works it cites.
Machine comprehension using match-lstm and answer pointer
S. Wang and J. Jiang · 2016
Later among the works it cites.
Multi-perspective context matching for machine comprehension
Z. Wang, H. Mi, W. Hamza, and R. Florian · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Visual storytelling
T.-H. K. Huang, F. Ferraro, N. Mostafazadeh, I. Misra, J. Devlin, A. Agrawal, R. Girshick, X. He, P. Kohli, D. Batra, et al · 2016
Cited alongside, same era.
Knowing when to look: Adaptive attention via A visual sentinel for image captioning
J. Lu, C. Xiong, D. Parikh, and R. Socher · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
J. Lu, J. Yang, D. Batra, and D. Parikh · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Cited alongside, same era.
Bidirectional attention flow for machine comprehension
M. Seo, A. Kembhavi, A. Farhadi, and H. Hajishirzi · 2016
Cited alongside, same era.
Later among the works it cites.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
C. Xiong, S. Merity, and R. Socher · 2016
Later among the works it cites.
Dynamic coattention networks for question answering
C. Xiong, V. Zhong, and R. Socher · 2016
Later among the works it cites.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2016
Later among the works it cites.
0swering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Delving deep into personal photo and video search
L. Jiang, L. Cao, Y. Kalantidis, S. Farfade, J. Tang, and A. G. Hauptmann · 2017
Closest in time.