Fetching the paper…
Reading the bibliography…
This paper addresses the problem of estimating and tracking human body keypoints in complex, multi-person video.
The hungarian method for the assignment problem
H. W. Kuhn · 1955
Earlier work this paper cites.
An algorithm for tracking multiple targets
D. B. Reid · 1979
Earlier work this paper cites.
Multi-target tracking using joint probabilistic data association
T. E. Fortman, Y. Bar-Shalom, and M. Scheffe · 1980
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Evaluating multiple object tracking performance: the CLEAR MOT metrics
K. Bernardin and R. Stiefelhagen · 2008
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Globally-optimal greedy algorithms for tracking a variable number of objects
H. Pirsiavash, D. Ramanan, and C. C. Fowlkes · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Articulated human detection with flexible mixtures of parts
Y. Yang and D. Ramanan · 2013
Earlier work this paper cites.
2D human pose estimation: New benchmark and state of the art analysis
M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Y.-W. Chao, Z. Wang, Y. He, J. Wang, and J. Deng · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. A. Hendricks, S. Guadarrama, S. V. M. Rohrbach, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Contextual action recognition with R*CNN
G. Gkioxari, R. Girshick, and J. Malik · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2016
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Quo vadis, action recognition? A new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Closest in time.
Spatiotemporal multiplier networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. P. Wildes · 2017
Closest in time.
Detect to track and track to detect
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2017
Closest in time.
Attentional pooling for action recognition
R. Girdhar and D. Ramanan · 2017
Closest in time.
ActionVLAD: Learning spatio-temporal aggregation for action classification
R. Girdhar, D. Ramanan, A. Gupta, J. Sivic, and B. Russell · 2017
Closest in time.
Mask R-CNN
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gurobi optimizer reference manual, 2016
I. Gurobi Optimization · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Deepercut: A deeper, stronger, and faster multi-person pose estimation model
E. Insafutdinov, L. Pishchulin, B. Andres, M. Andriluka, and B. Schiele · 2016
Cited alongside, same era.
T-CNN: Tubelets with convolutional neural networks for object detection from videos
K. Kang, H. Li, J. Yan, X. Zeng, B. Yang, T. Xiao, C. Zhang, Z. Wang, R. Wang, X. Wang, et al · 2016
Cited alongside, same era.
Object Detection from Video Tubelets with Convolutional Neural Networks
K. Kang, W. Ouyang, H. Li, and X. Wang · 2016
Cited alongside, same era.
Learning models for actions and person-object interactions with transfer to question answering
A. Mallya and S. Lazebnik · 2016
Cited alongside, same era.
R. Hou, C. Chen, and M. Shah · 2017
Closest in time.
Articulated multi-person tracking in the wild
E. Insafutdinov, M. Andriluka, L. Pishchulin, S. Tang, B. Andres, and B. Schiele · 2017
Closest in time.
PoseTrack: A benchmark for human pose estimation and tracking
U. Iqbal, A. Milan, M. Andriluka, E. Ensafutdinov, L. Pishchulin, J. Gall, and S. B · 2017
Closest in time.
PoseTrack dataset
U. Iqbal, A. Milan, M. Andriluka, E. Ensafutdinov, L. Pishchulin, J. Gall, and S. B · 2017
Closest in time.
Pose-track: Joint multi-person pose estimation and tracking
U. Iqbal, A. Milan, and J. Gall · 2017
Closest in time.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Closest in time.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Closest in time.
Online multi-target tracking using recurrent neural networks
A. Milan, S. H. Rezatofighi, A. R. Dick, I. D. Reid, and K. Schindler · 2017
Closest in time.
Towards accurate multi-person pose estimation in the wild
G. Papandreou, T. Zhu, N. Kanazawa, A. Toshev, J. Tompson, C. Bregler, and K. Murphy · 2017
Closest in time.
Tracking the untrackable: Learning to track multiple cues with long-term dependencies
A. Sadeghian, A. Alahi, and S. Savarese · 2017
Closest in time.
Thin-slicing network: A deep structured model for pose estimation in videos
J. Song, L. Wang, L. Van Gool, and O. Hilliges · 2017
Closest in time.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2017
Closest in time.