Fetching the paper…
Reading the bibliography…
Early action prediction deals with inferring the ongoing action from partially-observed videos, typically at the outset of the video.
Measures of the amount of ecologic association between species
Lee R Dice · 1945
Earlier work this paper cites.
Motor facilitation during action observation: a magnetic stimulation study
Luciano Fadiga, Leonardo Fogassi, Giovanni Pavesi, and Giacomo Rizzolatti · 1995
Earlier work this paper cites.
Action recognition in the premotor cortex
Vittorio Gallese, Luciano Fadiga, Leonardo Fogassi, and Giacomo Rizzolatti · 1996
Earlier work this paper cites.
Premotor cortex and the recognition of motor actions
Giacomo Rizzolatti, Luciano Fadiga, Vittorio Gallese, and Leonardo Fogassi · 1996
Earlier work this paper cites.
Hearing sounds, understanding actions: action representation in mirror neurons
Evelyne Kohler, Christian Keysers, M Alessandra Umilta, Leonardo Fogassi, Vittorio Gallese, and Giacomo Rizzolatti · 2002
Earlier work this paper cites.
Human activity prediction: Early recognition of ongoing activities from streaming videos
Michael S Ryoo · 2011
Earlier work this paper cites.
Modeling complex temporal composition of actionlets for activity prediction
Kang Li, Jie Hu, and Yun Fu · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Recognize human activities from partially observed videos
Yu Cao, Daniel Barrett, Andrei Barbu, Siddharth Narayanaswamy, Haonan Yu, Aaron Michaux, Yuewei Lin, Sven Dickinson, Jeffrey Mark Siskind, and Song Wang · 2013
Earlier work this paper cites.
A discriminative model with multiple temporal scales for action prediction
Yu Kong, Dmitry Kit, and Yun Fu · 2014
Earlier work this paper cites.
A hierarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese · 2014
Earlier work this paper cites.
Prediction of human activity by discovering temporal sequence patterns
Kang Li and Yun Fu · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Max-margin action prediction machine
Yu Kong and Yun Fu · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3D convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning activity progression in LSTMs for activity detection and early detection
Shugao Ma, Leonid Sigal, and Stan Sclaroff · 2016
Earlier work this paper cites.
NTU RGB+D: A large scale dataset for 3D human activity analysis
Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Deep sequential context networks for action prediction
Yu Kong, Zhiqiang Tao, and Yun Fu · 2017
Earlier work this paper cites.
Binary coding for partial action analysis with limited observation ratios
Jie Qin, Li Liu, Ling Shao, Bingbing Ni, Chen Chen, Fumin Shen, and Yunhong Wang · 2017
Earlier work this paper cites.
Encouraging LSTMs to anticipate actions very early
Mohammad Sadegh Aliakbarian, Fatemeh Sadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, and Lars Andersson · 2017
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from multi-view observation
Pierre Sermanet, Corey Lynch, Jasmine Hsu, and Sergey Levine · 2017
Cited alongside, same era.
A short note about kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Cited alongside, same era.
Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Early action prediction by soft regression
Jian-Fang Hu, Wei-Shi Zheng, Lianyang Ma, Gang Wang, Jianhuang Lai, and Jianguo Zhang · 2018
Cited alongside, same era.
Action prediction from videos via memorizing hard-to-predict samples
Yu Kong, Shangqian Gao, Bin Sun, and Yun Fu · 2018
Cited alongside, same era.
Adversarial action prediction networks
Rolling-unrolling LSTMs for action anticipation from first-person video
Antonino Furnari and Giovanni Maria Farinella · 2020
Later among the works it cites.
Confidence-guided self refinement for action prediction in untrimmed videos
Jingyi Hou, Xinxiao Wu, Ruiqi Wang, Jiebo Luo, and Yunde Jia · 2020
Later among the works it cites.
AR-Net: Adaptive frame resolution for efficient action recognition
Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris · 2020
Later among the works it cites.
A short note on the kinetics-700-2020 human action dataset
Lucas Smaira, João Carreira, Eric Noland, Ellen Clancy, Amy Wu, and Andrew Zisserman · 2020
Later among the works it cites.
Dynamic sampling networks for efficient action recognition in videos
Yin-Dong Zheng, Zhaoyang Liu, Tong Lu, and Limin Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu Kong, Zhiqiang Tao, and Yun Fu · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Temporal relational reasoning in videos
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba · 2018
Cited alongside, same era.
Action knowledge transfer for action prediction with partial videos
Yijun Cai, Haoxin Li, Jian-Fang Hu, and Wei-Shi Zheng · 2019
Cited alongside, same era.
Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution
Yunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng · 2019
Cited alongside, same era.
SlowFast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Cited alongside, same era.
ViViT: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Multiscale vision transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer · 2021
Later among the works it cites.
Anticipating human actions by correlating past with the future with jaccard similarity measures
Basura Fernando and Samitha Herath · 2021
Later among the works it cites.
Anticipative video transformer
Rohit Girdhar and Kristen Grauman · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira · 2021
Later among the works it cites.
MoViNets: Mobile video networks for efficient video recognition
Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong · 2021
Later among the works it cites.
TokenLearner: What can 8 learned tokens do for images and videos?
Michael S Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova · 2021
Later among the works it cites.
Adapool: Exponential adaptive pooling for information-retaining downsampling
Alexandros Stergiou and Ronald Poppe · 2021
Later among the works it cites.
Spatial–temporal relation reasoning for action prediction in videos
Xinxiao Wu, Ruiqi Wang, Jingyi Hou, Hanxi Lin, and Jiebo Luo · 2021
Later among the works it cites.
Anticipating future relations via graph growing for action prediction
Xinxiao Wu, Jianwei Zhao, and Ruiqi Wang · 2021
Later among the works it cites.
Long short-term transformer for online action detection
Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto · 2021
Later among the works it cites.
Multi-scale vision longformer: A new vision transformer for high-resolution image encoding
Pengchuan Zhang, Xiyang Dai, Jianwei Yang, Bin Xiao, Lu Yuan, Lei Zhang, and Jianfeng Gao · 2021
Later among the works it cites.
Vidtr: Video transformer without convolutions
Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe · 2021
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al · 2022
Closest in time.
ERA: Expert retrieval and assembly for early action prediction
Lin Geng Foo, Tianjiao Li, Hossein Rahmani, Qiuhong Ke, and Jun Liu · 2022
Closest in time.
Video Swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Closest in time.
MeMViT: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Closest in time.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Closest in time.