Fetching the paper…
Reading the bibliography…
Video understanding is a challenging problem with great impact on the abilities of autonomous agents working in the real-world.
A comparative analysis of selection schemes used in genetic algorithms
David E. Goldberg and Kalyanmoy Deb · 1991
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre · 2011
Earlier work this paper cites.
3d convolutional neural networks for human action recognition
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu · 2013
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
C3d: generic features for video analysis
Du Tran, Lubomir D Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2014
Earlier work this paper cites.
Every moment counts: Dense detailed labeling of actions in complex videos
Serena Yeung, Olga Russakovsky, Ning Jin, Mykhaylo Andriluka, Greg Mori, and Li Fei-Fei · 2015
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta · 2016
Earlier work this paper cites.
Leaving some stonesunturned: dynamic feature prioritization for activity detectionin streaming video
Yu-Chuan Su and Kristen Grauman · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin · 2017
Earlier work this paper cites.
Learnable pooling with context gating for video classification
Antoine Miech, Ivan Laptev, and Josef Sivic · 2017
Earlier work this paper cites.
Learning latent sub-events in activity videos using temporal attention filters
AJ Piergiovanni, Chenyou Fan, and Michael S Ryoo · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks
Zhaofan Qiu, Ting Yao, and Tao Mei · 2017
Earlier work this paper cites.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Yutaka Leon Suematsu Saurabh Saxena, Quoc Le, and Alex Kurakin · 2017
Earlier work this paper cites.
Asynchronous temporal fields for action recognition
Gunnar A Sigurdsson, Santosh Divvala, Ali Farhadi, and Abhinav Gupta · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc Le · 2017
Earlier work this paper cites.
Action search: Spotting actions in videos and its application to temporal action localization
Humam Alwassel, Fabian Caba Heilbron, and Bernard Ghanem · 2018
Earlier work this paper cites.
Massively parallel video networks
Joao Carreira, Viorica Patraucean, Laurent Mazare, and Andrew Zisserman · 2018
Earlier work this paper cites.
Multi-fiber networks for video recognition
Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng · 2018
Earlier work this paper cites.
Spatio-temporal channel correlation networks for action classification
Ali Diba, Mohsen Fayyaz, Vivek Sharma, M Mahdi Arzani, Rahman Yousefzadeh, Juergen Gall, and Luc Van Gool · 2018
Earlier work this paper cites.
Proxylessnas: Direct neural architecture search on target task and hardware
Song Han Han Cai, Ligeng Zhu · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Squeeze-and-excitation networks
Jie Hu, Li Shen, Samuel Albanie, Gang Sun, and Enhua Wu · 2018
Cited alongside, same era.
Motion feature network: Fixed motion filter for action recognition
Myunggi Lee, Seungeui Lee, Sungjoon Son, Gyutae Park, and Nojun Kwak · 2018
Cited alongside, same era.
Progressive neural architecture search
Chenxi Liu, Barret Zoph, Maxim Neumann, Jonathon Shlens, Wei Hua, Li-Jia Li, Li Fei-Fei, Alan Yuille, Jonathan Huang, and Kevin Murphy · 2018
Cited alongside, same era.
Moments in time dataset: one million videos for event understanding
Searching for mobilenetv3
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Vijay Vasudevan Ruoming Pang, Quoc V. Le, and Hartwig Adam · 2019
Closest in time.
Timeception for complex action recognition
Noureldien Hussein, Efstratios Gavves, and Arnold W.M. Smeulders · 2019
Closest in time.
Scsampler:sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Closest in time.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Closest in time.
DARTS: Differentiable architecture seach
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Closest in time.
Grouped spatial-temporalaggregation for efficient action recognition
Chenxu Luo and Alan L. Yuille · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfruend, Carl Vondrick, et al · 2018
Cited alongside, same era.
Efficient neural architecture search via parameter sharing
Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean · 2018
Cited alongside, same era.
Fine-grained activity recognition in baseball videos
AJ Piergiovanni and Michael S. Ryoo · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, , and L.-C. Chen · 2018
Cited alongside, same era.
Optical flow guided feature: a fast and robust motion representation for video action recognition
Shuyang Sun, Zhanghui Kuang, Lu Sheng, Wanli Ouyang, and Wei Zhang · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Evolving space-time neural architectures for videos
AJ Piergiovanni, Anelia Angelova, Alexander Toshev, and Michael S Ryoo · 2019
Closest in time.
Representation flow for action recognition
AJ Piergiovanni and Michael S Ryoo · 2019
Closest in time.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V. Le · 2019
Closest in time.
Mnasnet: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, and Quoc V Le · 2019
Closest in time.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Closest in time.
Fastdepth: Fast monocular depth estimation on embedded systems
Diana Wofk, Fangchang Ma, Tien-Ju Yang, Sertac Karaman, and Vivienne Sze · 2019
Closest in time.
Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search
Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer · 2019
Closest in time.
Scsampler:sampling salient clips from video for efficient action recognition
Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, and Shilei Wen · 2019
Closest in time.
Adaframe: Adaptive frame selection forfast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis · 2019
Closest in time.
Resource constrained neural network architecture search: Will a submodularity assumption help?
Yunyang Xiong, Ronak Mehta, and Vikas Singh · 2019
Closest in time.
Eena: Efficient evolution of neural architecture
Hui Zhu, Zhulin An, Chuanguang Yang, Kaiqiang Xu, Erhu Zhao, and Yongjun Xu · 2019
Closest in time.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Closest in time.
Assemblenet: Searching for multi-stream neural connectivity in video architectures
Michael S. Ryoo, AJ Piergiovanni, , Mingxing Tan, and Anelia Angelova · 2020
Closest in time.
Assemblenet++: Assembling modality representations via attention connections
Michael S. Ryoo, AJ Piergiovanni, Juhana Kangaspunta, and Anelia Angelova · 2020
Closest in time.