Fetching the paper…
Reading the bibliography…
Multi-modal learning, which focuses on utilizing various modalities to improve the performance of a model, is widely used in video recognition.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Marc’Aurelio Ranzato, and Tomas Mikolov · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Richard Socher, Milind Ganjoo, Christopher D Manning, and Andrew Ng · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri · 2015
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool · 2016
Earlier work this paper cites.
End-to-end learning of action detection from frame glimpses in videos
Serena Yeung, Olga Russakovsky, Greg Mori, and Li Fei-Fei · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Gated multimodal units for information fusion
John Arevalo, Thamar Solorio, Manuel Montes-y Gómez, and Fabio A González · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman · 2017
Earlier work this paper cites.
Spatially adaptive computation time for residual networks
Michael Figurnov, Maxwell D Collins, Yukun Zhu, Li Zhang, Jonathan Huang, Dmitry Vetrov, and Ruslan Salakhutdinov · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Exploiting feature and class relationships in video categorization with regularized deep neural networks
Yu-Gang Jiang, Zuxuan Wu, Jun Wang, Xiangyang Xue, and Shih-Fu Chang · 2017
Earlier work this paper cites.
Deciding how to decide: Dynamic routing in artificial neural networks
Mason McGill and Pietro Perona · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2017
Cited alongside, same era.
Watching a small portion could be as good as watching all: Towards efficient video classification
Hehe Fan, Zhongwen Xu, Linchao Zhu, Chenggang Yan, Jianjun Ge, and Yi Yang · 2018
Cited alongside, same era.
Dynamic zoom-in network for fast object detection in large images
Mingfei Gao, Ruichi Yu, Ang Li, Vlad I Morariu, and Larry S Davis · 2018
Cited alongside, same era.
Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet?
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh · 2018
Cited alongside, same era.
Efficient large-scale multi-modal classification
Douwe Kiela, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2018
Spottune: transfer learning through adaptive fine-tuning
Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris · 2019
Later among the works it cites.
Channel gating neural networks
Weizhe Hua, Yuan Zhou, Christopher M De Sa, Zhiru Zhang, and G Edward Suh · 2019
Later among the works it cites.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen · 2019
Later among the works it cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2019
Later among the works it cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Later among the works it cites.
Autofocus: Efficient multi-scale inference
Mahyar Najibi, Bharat Singh, and Larry S Davis · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Motion feature network: Fixed motion filter for action recognition
Myunggi Lee, Seungeui Lee, Sungjoon Son, Gyutae Park, and Nojun Kwak · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Optical flow guided feature: A fast and robust motion representation for video action recognition
Shuyang Sun, Zhanghui Kuang, Lu Sheng, Wanli Ouyang, and Wei Zhang · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Convolutional networks with adaptive inference graphs
Andreas Veit and Serge Belongie · 2018
Cited alongside, same era.
Later among the works it cites.
Mfas: Multimodal fusion architecture search
Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux, Moez Baccouche, and Frédéric Jurie · 2019
Later among the works it cites.
AJ Piergiovanni, Anelia Angelova, and Michael S Ryoo · 2019
Later among the works it cites.
Adashare: Learning what to share for efficient deep multi-task learning
Ximeng Sun, Rameswar Panda, and Rogerio Feris · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
Video classification with channel-separated convolutional networks
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli · 2019
Later among the works it cites.
Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search
Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer · 2019
Later among the works it cites.
Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition
Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, and Shilei Wen · 2019
Later among the works it cites.
Liteeval: A coarse-to-fine framework for resource efficient video recognition
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S Davis · 2019
Later among the works it cites.
Adaframe: Adaptive frame selection for fast video recognition
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis · 2019
Later among the works it cites.
X3d: Expanding architectures for efficient video recognition
Christoph Feichtenhofer · 2020
Later among the works it cites.
Listen to look: Action recognition by previewing audio
Gao, Ruohan and Oh, Tae-Hyun, and Grauman, Kristen and Torresani, Lorenzo · 2020
Later among the works it cites.
Timegate: Conditional gating of segments in long-range activities
Noureldien Hussein, Mihir Jain, and Babak Ehteshami Bejnordi · 2020
Later among the works it cites.
Ar-net: Adaptive frame resolution for efficient action recognition
Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris · 2020
Later among the works it cites.
What makes training multi-modal networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.