2020

RareAct: A video dataset of unusual interactions

Miech, Antoine, Alayrac, Jean-Baptiste, Laptev, Ivan et al.

Understand

This paper introduces a manually annotated video dataset of unusual actions, namely RareAct, including actions such as "blend phone", "cut keyboard" and "microwave shoes".

  • RareAct aims at evaluating the zero-shot and few-shot compositionality of action recognition models for unlikely compositions of common action verbs and object nouns.
  • It contains 122 different actions which were obtained by combining verbs and nouns rarely co-occurring together in the large-scale textual corpus from HowTo100M, but that frequently appear separately.
  • We provide benchmarks using a state-of-the-art HowTo100M pretrained video and text model and show that zero-shot and few-shot compositionality of actions remains a challenging and unsolved task.

Built on

  • Context models and out-of-context objects

    Myung Jin Choi, Antonio Torralba, and Alan S. Willsky · 2012

    Earlier work this paper cites.

  • Expecting the unexpected: Training detectors for unusual pedestrians with adversarial imposters

    Shiyu Huang and Deva Ramanan · 2017

    Earlier work this paper cites.

  • The kinetics human action video dataset

    Original

    Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017

    Earlier work this paper cites.

Similar

  • Weakly-supervised learning of visual relations

    Julia Peyre, Ivan Laptev, Cordelia Schmid, and Josef Sivic · 2017

    Cited alongside, same era.

  • Scaling egocentric vision: The epic-kitchens dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2018

    Cited alongside, same era.

  • Ava: A video dataset of spatio-temporally localized atomic visual actions

    Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, et al · 2018

    Cited alongside, same era.

Then

  • Howto100M: Learning a text-video embedding by watching hundred million narrated video clips

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019

    Later among the works it cites.

  • Oops! predicting unintentional action in video

    Dave Epstein, Boyuan Chen, and Carl Vondrick · 2020

    Closest in time.

  • End-to-end learning of visual representations from uncurated instructional videos

    Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…