Fetching the paper…
Reading the bibliography…
Video object detection is more challenging compared to image object detection.
Determining optical flow
B. K. Horn and B. G. Schunck · 1981
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
L. Wiskott and T. J. Sejnowski · 2002
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert · 2004
Earlier work this paper cites.
Slow feature analysis for human action recognition
Z. Zhang and D. Tao · 2012
Earlier work this paper cites.
Deep learning of invariant features via simulated fixations in video
W. Zou, S. Zhu, K. Yu, and A. Y. Ng · 2012
Earlier work this paper cites.
Deepflow: Large displacement optical flow with deep matching
P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Dl-sfa: deeply-learned slow feature analysis for action recognition
L. Sun, K. Jia, T.-H. Chan, Y. Fang, G. Wang, and S. Yan · 2014
Earlier work this paper cites.
Delving deeper into convolutional networks for learning video representations
N. Ballas, L. Yao, C. Pal, and A. Courville · 2015
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. van der Smagt, D. Cremers, and T. Brox · 2015
Earlier work this paper cites.
Fast r-cnn
R. Girshick · 2015
Cited alongside, same era.
Spatial transformer networks
M. Jaderberg, K. Simonyan, A. Zisserman, et al · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Epicflow: Edge-preserving interpolation of correspondences for optical flow
J. Revaud, P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Cited alongside, same era.
Action recognition using visual attention
S. Sharma, R. Kiros, and R. Salakhutdinov · 2015
Slow and steady feature analysis: higher order temporal coherence in video
D. Jayaraman and K. Grauman · 2016
Later among the works it cites.
A. Kar, N. Rai, K. Sikka, and G. Sharma · 2016
Later among the works it cites.
Multi-class multi-object tracking using changing point detection
B. Lee, E. Erdenee, S. Jin, M. Y. Nam, Y. G. Jung, and P. K. Rhee · 2016
Later among the works it cites.
Clockwork convnets for video semantic segmentation
E. Shelhamer, K. Rakelly, J. Hoffman, and T. Darrell · 2016
Later among the works it cites.
Deep feature flow for video recognition
X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
R-fcn: Object detection via region-based fully convolutional networks
J. Dai, Y. Li, K. He, and J. Sun · 2016
Cited alongside, same era.
Seq-nms for video object detection
W. Han, P. Khorrami, T. L. Paine, P. Ramachandran, M. Babaeizadeh, H. Shi, J. Li, S. Yan, and T. S. Huang · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Flownet 2.0: Evolution of optical flow estimation with deep networks
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox · 2016
Cited alongside, same era.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei · 2017
Closest in time.
Semantic video cnns through representation warping
R. Gadde, V. Jampani, and P. V. Gehler · 2017
Closest in time.
T-cnn: Tubelets with convolutional neural networks for object detection from videos
K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao, C. Zhang, Z. Wang, R. Wang, X. Wang, et al · 2017
Closest in time.
Videolstm convolves, attends and flows for action recognition
Z. Li, K. Gavrilyuk, E. Gavves, M. Jain, and C. G. Snoek · 2017
Closest in time.
Quality aware network for set to set recognition
Y. Liu, J. Yan, and W. Ouyang · 2017
Closest in time.
Flow-guided feature aggregation for video object detection
X. Zhu, Y. Wang, J. Dai, L. Yuan, and Y. Wei · 2017
Closest in time.