Fetching the paper…
Reading the bibliography…
Existing sketch-analysis work studies sketches depicting static objects or scenes.
G. A. Miller, “The magical number seven, plus or minus two: Some limits on our capacity for processing information.”
1956
Earlier work this paper cites.
E. Tulving and D. Murray, “Elements of episodic memory,”
1985
Earlier work this paper cites.
Y.-P. Tan, S. R. Kulkarni, and P. J. Ramadge, “A framework for measuring video similarity and its application to video query by example,” in
1999
Earlier work this paper cites.
S. Andrews, I. Tsochantaridis, and T. Hofmann, “Support vector machines for multiple-instance learning,” in
2003
Earlier work this paper cites.
B. Coskun, B. Sankur, and N. Memon, “Spatio–temporal transform based video hashing,”
2006
Earlier work this paper cites.
C. Zach, T. Pock, and H. Bischof, “A duality based approach for realtime
2007
Earlier work this paper cites.
J. P. Collomosse, G. McNeill, and Y. Qian, “Storyboard sketches for content based video retrieval,” in
2009
Earlier work this paper cites.
R. Hu and J. Collomosse, “Motion-sketch based video retrieval using a trellis levenshtein distance,” in
2010
Earlier work this paper cites.
R. Hu, S. James, and J. Collomosse, “Annotated free-hand sketches for video retrieval using object semantics and motion,” in
2012
Earlier work this paper cites.
M. Li and V. Monga, “Robust video hashing via multilinear subspace projections,”
2012
Earlier work this paper cites.
M. Eitz, J. Hays, and M. Alexa, “How do humans sketch objects?”
2012
Earlier work this paper cites.
R. Hu, S. James, T. Wang, and J. Collomosse, “Markov random fields for sketch based video retrieval,” in
2013
Earlier work this paper cites.
G. Ye, D. Liu, J. Wang, and S.-F. Chang, “Large-scale video hashing via structure learning,” in
2013
Earlier work this paper cites.
S. James and J. Collomosse, “Interactive video asset retrieval using sketched queries,” in
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in
2014
Earlier work this paper cites.
Y. Li, R. Wang, Z. Huang, S. Shan, and X. Chen, “Face video retrieval with image query via hashing across euclidean space and riemannian manifold,” in
2015
Earlier work this paper cites.
R. Xu, C. Xiong, W. Chen, and J. J. Corso, “Jointly modeling deep video and compositional text to bridge vision and language in a unified framework.” in
2015
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in
2015
Earlier work this paper cites.
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in
2015
Cited alongside, same era.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in
2015
Cited alongside, same era.
Y. Song, J. Vallmitjana, A. Stent, and A. Jaimes, “Tvsum: Summarizing web videos using titles,” in
2015
Cited alongside, same era.
W.-S. Chu, Y. Song, and A. Jaimes, “Video co-summarization: Video summarization by visual co-occurrence,” in
2015
Cited alongside, same era.
Q. Yu, F. Liu, Y.-Z. Song, T. Xiang, T. M. Hospedales, and C.-C. Loy, “Sketch me that shoe,” in
2016
Cited alongside, same era.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in
2018
Later among the works it cites.
K. Hara, H. Kataoka, and Y. Satoh, “Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?” in
2018
Later among the works it cites.
K. Liu, W. Liu, C. Gan, M. Tan, and H. Ma, “T-c3d: Temporal convolutional 3d network for real-time action recognition,” in
2018
Later among the works it cites.
F. Sung, Y. Yang, L. Zhang, T. Xiang, P. H. Torr, and T. M. Hospedales, “Learning to compare: Relation network for few-shot learning,” in
2018
Later among the works it cites.
W. Xie, L. Shen, and A. Zisserman, “Comparator networks,” in
2018
Later among the works it cites.
A. Araujo and B. Girod, “Large-scale video retrieval using image queries,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Y. Ye, Y. Lu, and H. Jiang, “Human’s scene sketch understanding,” in
2016
Cited alongside, same era.
P. Sangkloy, N. Burnell, C. Ham, and J. Hays, “The sketchy database: learning to retrieve badly drawn bunnies,”
2016
Cited alongside, same era.
B. Singh, T. Marks, M. Jones, O. Tuzel, and M. Shao, “A multi-stream bi-directional recurrent neural network for fine-grained action detection,” in
2016
Cited alongside, same era.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks: Towards good practices for deep action recognition,” in
2016
Cited alongside, same era.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in
2016
Cited alongside, same era.
Z. Qiu, T. Yao, and T. Mei, “Learning spatio-temporal representation with pseudo-3d residual networks,” in
2017
Cited alongside, same era.
2018
Later among the works it cites.
N. C. Mithun, J. Li, F. Metze, and A. K. Roy-Chowdhury, “Learning joint embedding with multimodal cues for cross-modal video-text retrieval,” in
2018
Later among the works it cites.
Z. Chen, J. Lu, J. Feng, and J. Zhou, “Nonlinear structural hashing for scalable video search,”
2018
Later among the works it cites.
J. Song, H. Zhang, X. Li, L. Gao, M. Wang, and R. Hong, “Self-supervised video hashing with hierarchical binary auto-encoder,”
2018
Later among the works it cites.
L. Wang, W. Li, W. Li, and L. Van Gool, “Appearance-and-relation networks for video classification,” in
2018
Later among the works it cites.
A. Diba, M. Fayyaz, V. Sharma, M. Mahdi Arzani, R. Yousefzadeh, J. Gall, and L. Van Gool, “Spatio-temporal channel correlation networks for action classification,” in
2018
Later among the works it cites.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in
2018
Later among the works it cites.
Y. Chen, Y. Kalantidis, J. Li, S. Yan, and J. Feng, “Multi-fiber networks for video recognition,” in
2018
Later among the works it cites.
T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, G. Liu, A. Tao, J. Kautz, and B. Catanzaro, “Video-to-video synthesis,” in
2018
Later among the works it cites.
J. Collomosse, T. Bui, and H. Jin, “Livesketch: Query perturbations for guided sketch-based visual search,” in
2019
Later among the works it cites.
Y. Xie, P. Xu, and Z. Ma, “Deep zero-shot learning for scene sketch,”
2019
Later among the works it cites.
P. Xu, “Deep learning for free-hand sketch: A survey,”
2020
Closest in time.