2020

The AVA-Kinetics Localized Human Actions Video Dataset

Li, Ang, Thotakuri, Meghana, Ross, David A. et al.

Understand

This paper describes the AVA-Kinetics localized human actions video dataset.

  • The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these new AVA annotated Kinetics clips.
  • The dataset contains over 230k clips annotated with the 80 AVA action classes for each of the humans in key-frames.
  • We describe the annotation process and provide statistics about the new dataset.

Built on

  • ImageNet: A Large-Scale Hierarchical Image Database

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009

    Earlier work this paper cites.

  • Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014

    Earlier work this paper cites.

  • Faster r-cnn: Towards real-time object detection with region proposal networks, 2015

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015

    Earlier work this paper cites.

  • Hollywood in homes: Crowdsourcing data collection for activity understanding

    Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016

    Earlier work this paper cites.

Similar

  • The kinetics human action video dataset

    Original

    Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman · 2017

    Cited alongside, same era.

  • Ava: A video dataset of spatio-temporally localized atomic visual actions

    Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and Jitendra Malik · 2018

    Cited alongside, same era.

  • A short note on the kinetics-700 human action dataset

    Original

    Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman · 2019

    Cited alongside, same era.

Then

  • Slowfast networks for video recognition

    Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019

    Later among the works it cites.

  • Video action transformer network

    Rohit Girdhar, Joao Carreira, Carl Doersch, and Andrew Zisserman · 2019

    Later among the works it cites.

  • Objects as points

    Original

    Xingyi Zhou, Dequan Wang, and Philipp Krähenbühl · 2019

    Later among the works it cites.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…