2019

CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Yi, Kexin, Gan, Chuang, Li, Yunzhu et al.

Understand

The ability to reason about temporal and causal events from videos lies at the core of human intelligence.

  • Most video reasoning benchmarks, however, focus on pattern recognition from complex visual and language input, instead of on causal structure.
  • We study the complementary problem, exploring the temporal and causal structures behind videos of objects with simple visual appearance.
  • To this end, we introduce the CoLlision Events for Video REpresentation and Reasoning (CLEVRER), a diagnostic video dataset for systematic evaluation of computational models on a wide range of reasoning tasks.

Reading the bibliography…