2020

Diverse Temporal Aggregation and Depthwise Spatiotemporal Factorization for Efficient Video Classification

Lee, Youngwan, Kim, Hyung-Il, Yun, Kimin et al.

Understand

Video classification researches that have recently attracted attention are the fields of temporal modeling and 3D efficient architecture.

  • However, the temporal modeling methods are not efficient or the 3D efficient architecture is less interested in temporal modeling.
  • For bridging the gap between them, we propose an efficient temporal modeling 3D architecture, called VoV3D, that consists of a temporal one-shot aggregation (T-OSA) module and depthwise factorized component, D(2+1)D.
  • The T-OSA is devised to build a feature hierarchy by aggregating temporal features with different temporal receptive fields.

Reading the bibliography…