2018

Videos as Space-Time Region Graphs

Wang, Xiaolong, Gupta, Abhinav

Understand

How do humans recognize the action "opening a book" ? We argue that there are two important cues: modeling temporal shape dynamics and modeling functional relationships between humans and objects.

  • In this paper, we propose to represent videos as space-time region graphs which capture these two important cues.
  • Our graph nodes are defined by the object region proposals from different frames in a long range video.
  • These nodes are connected by two types of relations: (i) similarity relations capturing the long range dependencies between correlated objects and (ii) spatial-temporal relations capturing the interactions between nearby objects.

Reading the bibliography…