Fetching the paper…
Reading the bibliography…
We describe an extension of the DeepMind Kinetics human action dataset from 600 classes to 700 classes, where for each class there are at least 600 video clips from different YouTube videos.
Y. Bian, C. Gan, X. Liu, F. Li, X. Long, Y. Li, H. Qi, J. Zhou, S. Wen, and Y. Lin · 2017
Earlier work this paper cites.
Quo Vadis, Action Recognition? New Models and the Kinetics Dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
AVA: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Cited alongside, same era.
A short note about Kinetics-600
J. Carreira, E. Noland, A. Banki-Horvath, C. Hillier, and A. Zisserman · 2018
Cited alongside, same era.
Scaling Egocentric Vision: The EPIC-KITCHENS Dataset
D. Damen, H. Doughty, G. Maria Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2018
Later among the works it cites.
Exploiting spatial-temporal modelling and multi-modal fusion for human action recognition
D. He, F. Li, Q. Zhao, X. Long, Y. Fu, and S. Wen · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…