Fetching the paper…
Reading the bibliography…
Videos are multimodal in nature.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 1998 · 1998
Earlier work this paper cites.
Concept-driven multi-modality fusion for video search
Xiao-Yong Wei, Yu-Gang Jiang, and Chong-Wah Ngo. 2011 · 2011
Earlier work this paper cites.
Action Recognition with Improved Trajectories. In ICCV
Heng Wang and Cordelia Schmid. 2013 · 2013
Earlier work this paper cites.
Discovering Joint Audio-Visual Codewords for Video Event Detection
I-Hong Jhuo, Guangnan Ye, Shenghua Gao, Dong Liu, Yu-Gang Jiang, D. T. Lee, and Shih-Fu Chang. 2014 · 2014
Earlier work this paper cites.
Two-Stream Convolutional Networks for Action Recognition in Videos. In NIPS
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding. In CVPR
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Super fast event recognition in internet videos
Yu-Gang Jiang, Qi Dai, Tao Mei, Yong Rui, and Shih-Fu Chang. 2015 · 2015
Earlier work this paper cites.
Beyond Short Snippets: Deep Networks for Video Classification. In CVPR
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015 · 2015
Earlier work this paper cites.
C3D: Generic Features for Video Analysis. In ICCV
Du Tran, Lubomir D Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Evaluating Two-Stream CNN for Video Classification. In ACM ICMR
Hao Ye, Zuxuan Wu, Rui-Wei Zhao, Xi Wang, Yu-Gang Jiang, and Xiangyang Xue. 2015 · 2015
Earlier work this paper cites.
Query-adaptive late fusion for image search and person re-identification. In CVPR
Liang Zheng, Shengjin Wang, Lu Tian, Fei He, Ziqiong Liu, and Qi Tian. 2015 · 2015
Earlier work this paper cites.
Convolutional Two-Stream Network Fusion for Video Action Recognition. In CVPR
C. Feichtenhofer, A. Pinz, and A. Zisserman. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
Zhaofan Qiu, Ting Yao, and Tao Mei. 2016 · 2016
Earlier work this paper cites.
Multi-Stream Multi-Class Fusion of Deep Networks for Video Classification. In ACM Multimedia
Zuxuan Wu, Yu-Gang Jiang, Xi Wang, Hao Ye, and Xiangyang Xue. 2016 · 2016
Earlier work this paper cites.
Multilayer and multimodal fusion of deep neural networks for video classification. In ACM Multimedia
Xiaodong Yang, Pavlo Molchanov, and Jan Kautz. 2016 · 2016
Earlier work this paper cites.
Real-time action recognition with enhanced motion vector CNNs. In CVPR
Bowen Zhang, Limin Wang, Zhe Wang, Yu Qiao, and Hanli Wang. 2016 · 2016
Earlier work this paper cites.
Look, listen and learn. In ICCV
Relja Arandjelovic and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Adaptive neural networks for fast test-time prediction. In ICML
Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama. 2017 · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events. In ICASSP
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax. In ICLR
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables. In ICLR
Chris J Maddison, Andriy Mnih, and Yee Whye Teh. 2017 · 2017
Cited alongside, same era.
Deep quantization: Encoding convolutional activations with deep generative model. In CVPR
Zhaofan Qiu, Ting Yao, and Tao Mei. 2017 · 2017
Cited alongside, same era.
Modeling multimodal clues in a hybrid deep learning framework for video classification
Yu-Gang Jiang, Zuxuan Wu, Jinhui Tang, Zechao Li, Xiangyang Xue, and Shih-Fu Chang. 2018a · 2018
Cited alongside, same era.
Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
Yu-Gang Jiang, Zuxuan Wu, Jun Wang, Xiangyang Xue, and Shih-Fu Chang. 2018b · 2018
Cited alongside, same era.
Pivot correlational neural network for multimodal video categorization. In ECCV
Sunghun Kang, Junyeong Kim, Hyunsoo Choi, Sungjin Kim, and Chang D Yoo. 2018 · 2018
Cited alongside, same era.
Dense dilated network for video action recognition
Baohan Xu, Hao Ye, Yingbin Zheng, Heng Wang, Tianyu Luwang, and Yu-Gang Jiang. 2019 · 2019
Later among the works it cites.
Visual content recognition by exploiting semantic feature map with attention and multi-task learning
Rui-Wei Zhao, Qi Zhang, Zuxuan Wu, Jianguo Li, and Yu-Gang Jiang. 2019 · 2019
Later among the works it cites.
Vision-infused deep audio inpainting. In ICCV
Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Later among the works it cites.
End-to-end object detection with transformers. In ECCV
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Later among the works it cites.
X3d: Expanding architectures for efficient video recognition. In CVPR
Christoph Feichtenhofer. 2020 · 2020
Later among the works it cites.
Exploring deep learning for view-based 3D model retrieval
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Videolstm convolves, attends and flows for action recognition
Zhenyang Li, Kirill Gavrilyuk, Efstratios Gavves, Mihir Jain, and Cees GM Snoek. 2018 · 2018
Cited alongside, same era.
Multimodal keyless attention fusion for video classification. In AAAI
Xiang Long, Chuang Gan, Gerard Melo, Xiao Liu, Yandong Li, Fu Li, and Shilei Wen. 2018 · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features. In ECCV
Andrew Owens and Alexei A Efros. 2018 · 2018
Cited alongside, same era.
MobileNetV2: Inverted Residuals and Linear Bottlenecks. In CVPR
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. 2018 · 2018
Cited alongside, same era.
Audio-visual event localization in unconstrained videos. In ECCV
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition. In CVPR
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018 · 2018
Cited alongside, same era.
Activity recognition using temporal optical flow convolutional features and multilayer LSTM
Amin Ullah, Khan Muhammad, Javier Del Ser, Sung Wook Baik, and Victor Hugo C de Albuquerque. 2018 · 2018
Cited alongside, same era.
Zan Gao, Yinming Li, and Shaohua Wan. 2020a · 2020
Later among the works it cites.
Ar-net: Adaptive frame resolution for efficient action recognition. In ECCV
Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris. 2020 · 2020
Later among the works it cites.
Towards More Explainability: Concept Knowledge Mining Network for Event Recognition. In ACM Multimedia
Zhaobo Qi, Shuhui Wang, Chi Su, Li Su, Qingming Huang, and Qi Tian. 2020 · 2020
Later among the works it cites.
Modality compensation network: Cross-modal adaptation for action recognition
Sijie Song, Jiaying Liu, Yanghao Li, and Zongming Guo. 2020 · 2020
Later among the works it cites.
Learning when and where to zoom with deep reinforcement learning. In CVPR
Burak Uzkent and Stefano Ermon. 2020 · 2020
Later among the works it cites.
A Dynamic Frame Selection Framework for Fast Video Recognition
Zuxuan Wu, Hengduo Li, Caiming Xiong, Yu-Gang Jiang, and Larry Steven Davis. 2020 · 2020
Later among the works it cites.
Audiovisual slowfast networks for video recognition
Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer. 2020 · 2020
Later among the works it cites.
STA-CNN: convolutional spatial-temporal attention learning for action recognition
Hao Yang, Chunfeng Yuan, Li Zhang, Yunda Sun, Weiming Hu, and Stephen J Maybank. 2020b · 2020
Later among the works it cites.
Faster recurrent networks for efficient video classification. In AAAI
Linchao Zhu, Du Tran, Laura Sevilla-Lara, Yi Yang, Matt Feiszli, and Heng Wang. 2020 · 2020
Later among the works it cites.
Conquer: Contextual query-aware ranking for video corpus moment retrieval. In ACM Multimedia
Zhijian Hou, Chong-Wah Ngo, and Wing Kwong Chan. 2021 · 2021
Closest in time.
Selective Feature Compression for Efficient Activity Recognition Inference
Chunhui Liu, Xinyu Li, Hao Chen, Davide Modolo, and Joseph Tighe. 2021 · 2021
Closest in time.
AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition. In ICLR
Yue Meng, Rameswar Panda, Chung-Ching Lin, Prasanna Sattigeri, Leonid Karlinsky, Kate Saenko, Aude Oliva, and Rogerio Feris. 2021 · 2021
Closest in time.
VA-RED 2
Bowen Pan, Rameswar Panda, Camilo Fosco, Chung-Ching Lin, Alex Andonian, Yue Meng, Kate Saenko, Aude Oliva, and Rogerio Feris. 2021 · 2021
Closest in time.
Multi-level Temporal Dilated Dense Prediction for Action Recognition
Jinpeng Wang, Yiqi Lin, Manlin Zhang, Yuan Gao, and Andy J Ma. 2021 · 2021
Closest in time.
Semi-supervised vision transformers
Zejia Weng, Xitong Yang, Ang Li, Zuxuan Wu, and Yu-Gang Jiang. 2021 · 2021
Closest in time.
MMSUM digital twins: a multi-view multi-modality summarization framework for sporting events
Samah Aloufi and Abdulmotaleb El Saddik. 2022 · 2022
Closest in time.
Local Correlation Ensemble with GCN based on Attention Features for Cross-domain Person Re-ID
Yue Zhang, Fanghui Zhang, Yi Jin, Yigang Cen, Viacheslav Voronin, and Shaohua Wan. 2022 · 2022
Closest in time.
Clustering Matters: Sphere Feature for Fully Unsupervised Person Re-identification
Yi Zheng, Yong Zhou, Jiaqi Zhao, Ying Chen, Rui Yao, Bing Liu, and Abdulmotaleb El Saddik. 2022 · 2022
Closest in time.