Dynamonet: Dynamic action and motion network
Ali Diba, Vivek Sharma, Luc Van Gool, and Rainer Stiefelhagen · 2019
Later among the works it cites.
More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation
Quanfu Fan, Chun-Fu Richard Chen, Hilde Kuehne, Marco Pistoia, and David Cox · 2019
Later among the works it cites.
Listen to look: Action recognition by previewing audio
Original
Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani · 2019
Later among the works it cites.
Spottune: transfer learning through adaptive fine-tuning
Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris · 2019
Later among the works it cites.
Channel gating neural networks
Weizhe Hua, Yuan Zhou, Christopher M De Sa, Zhiru Zhang, and G Edward Suh · 2019
Later among the works it cites.
You only watch once: A unified cnn architecture for real-time spatiotemporal action localization
Okan Köpüklü, Xiangyu Wei, and Gerhard Rigoll · 2019
Later among the works it cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Later among the works it cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
Moments in time dataset: one million videos for event understanding
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Original
Mingxing Tan and Quoc V Le · 2019
Later among the works it cites.
Mnasnet: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le · 2019
Later among the works it cites.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Later among the works it cites.
Adaframe: Adaptive frame selection for fast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis · 2019
Later among the works it cites.
Rubiksnet: Learnable 3d-shift for efficient video action recognition
Linxi Fan, Shyamal Buch, Guanzhi Wang, Ryan Cao, Yuke Zhu, Juan Carlos Niebles, and Li Fei-Fei · 2020
Later among the works it cites.
X3d: Expanding architectures for efficient video recognition
Original
Christoph Feichtenhofer · 2020
Later among the works it cites.
Directional temporal modeling for action recognition
Xinyu Li, Bing Shuai, and Joseph Tighe · 2020
Later among the works it cites.
Ar-net: Adaptive frame resolution for efficient action recognition
Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris · 2020
Later among the works it cites.
Video modeling with correlation networks
Heng Wang, Du Tran, Lorenzo Torresani, and Matt Feiszli · 2020
Later among the works it cites.
Temporal pyramid network for action recognition
Ceyuan Yang, Yinghao Xu, Jianping Shi, Bo Dai, and Bolei Zhou · 2020
Later among the works it cites.
Adafuse: Adaptive temporal fusion network for efficient action recognition
Yue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri, Leonid Karlinsky, Kate Saenko, Aude Oliva, and Rogerio Feris · 2021
Closest in time.