Fetching the paper…
Reading the bibliography…
Extending state-of-the-art object detectors from image to video is challenging.
Determining optical flow
B. K. Horn and B. G. Schunck · 1981
Earlier work this paper cites.
High accuracy optical flow estimation based on a theory for warping
T. Brox, A. Bruhn, N. Papenberg, and J. Weickert · 2004
Earlier work this paper cites.
A survey on variational optic flow methods for small displacements
J. Weickert, A. Bruhn, T. Brox, and N. Papenberg · 2006
Earlier work this paper cites.
Large displacement optical flow: descriptor matching in variational motion estimation
T. Brox and J. Malik · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Deepflow: Large displacement optical flow with deep matching
P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
R. Girshick, J. Donahue, T. Darrell, and J. Malik · 2014
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Flownet: Learning optical flow with convolutional networks
A. Dosovitskiy, P. Fischer, E. Ilg, P. Hausser, C. Hazirbas, V. Golkov, P. v.d. Smagt, D. Cremers, and T. Brox · 2015
Earlier work this paper cites.
Fast r-cnn
R. Girshick · 2015
Earlier work this paper cites.
Visual tracking with fully convolutional networks
W. Lijun, O. Wanli, W. Xiaogang, and L. Huchuan · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Epicflow: Edge-preserving interpolation of correspondences for optical flow
J. Revaud, P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
A. M. Rush, S. Chopra, and J. Weston · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. Berg, and F.-F. Li · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Human action recognition using factorized spatio-temporal convolutional networks
L. Sun, K. Jia, D.-Y. Yeung, and B. E. Shi · 2015
Cited alongside, same era.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Multi-class multi-object tracking using changing point detection
B. Lee, E. Erdenee, S. Jin, M. Y. Nam, Y. G. Jung, and P. K. Rhee · 2016
Later among the works it cites.
Videolstm convolves, attends and flows for action recognition
Z. Li, E. Gavves, M. Jain, and C. G. Snoek · 2016
Later among the works it cites.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Later among the works it cites.
Optical flow estimation using a spatial pyramid network
A. Ranjan and M. J. Black · 2016
Later among the works it cites.
Action recognition using visual attention
S. Sharma, R. Kiros, and R. Salakhutdinov · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Cited alongside, same era.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Cited alongside, same era.
Delving deeper into convolutional networks for learning video representations
N. Ballas, L. Yao, C. Pal, and A. Courville · 2016
Cited alongside, same era.
R-fcn: Object detection via region-based fully convolutional networks
J. Dai, Y. Li, K. He, and J. Sun · 2016
Cited alongside, same era.
Stfcn: Spatio-temporal fcn for semantic video segmentation
M. Fayyaz, M. Hajizadeh Saffar, M. Sabokrou, M. Fathy, R. Klette, and F. Huang · 2016
Cited alongside, same era.
Seq-nms for video object detection
W. Han, P. Khorrami, T. Le Paine, P. Ramachandran, M. Babaeizadeh, H. Shi, J. Li, S. Yan, and T. S. Huang · 2016
Cited alongside, same era.
M. Siam, S. Valipour, M. Jagersand, and N. Ray · 2016
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi · 2016
Later among the works it cites.
Deep end2end voxel2voxel prediction
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2016
Later among the works it cites.
Efficient object detection from videos
J. Yang, H. Shuai, Z. Yu, R. Fan, Q. Ma, Q. Liu, and J. Deng · 2016
Later among the works it cites.
On the stability of video detection and tracking
H. Zhang and N. Wang · 2016
Later among the works it cites.
Deformable convolutional networks
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei · 2017
Closest in time.
Flownet 2.0: Evolution of optical flow estimation with deep networks
E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox · 2017
Closest in time.
Adascan: Adaptive scan pooling in deep convolutional neural networks for human action recognition in videos
A. Kar, N. Rai, K. Sikka, and G. Sharma · 2017
Closest in time.
Cosine normalization: Using cosine similarity instead of dot product in neural networks
C. Luo, J. Zhan, L. Wang, and Q. Yang · 2017
Closest in time.
Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video
E. Real, J. Shlens, S. Mazzocchi, X. Pan, and V. Vanhoucke · 2017
Closest in time.
Deep feature flow for video recognition
X. Zhu, Y. Xiong, J. Dai, L. Yuan, and Y. Wei · 2017
Closest in time.