Fetching the paper…
Reading the bibliography…
In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical.
Eco: Efficient convolutional network for online video understanding
M. Zolfaghari, K. Singh, and T. Brox · 1912
Earlier work this paper cites.
What in the world do we hear?: An ecological approach to auditory event perception
W. W. Gaver · 1993
Earlier work this paper cites.
Space-time interest points
I. Laptev and T. Lindeberg · 2003
Earlier work this paper cites.
An efficient dense and scale-invariant spatio-temporal interest point detector
G. Willems, T. Tuytelaars, and L. Van Gool · 2008
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Y. J. Lee, J. Ghosh, and K. Grauman · 2012
Earlier work this paper cites.
Discovering discriminative action parts from mid-level video representations
M. Raptis, I. Kokkinos, and S. Soatto · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
K. Soomro, A. R. Zamir, and M. Shah · 2012
Earlier work this paper cites.
Representing videos using mid-level discriminative patches
A. Jain, A. Gupta, M. Rodriguez, and L. S. Davis · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
Motionlets: Mid-level 3d parts for human motion recognition
L. Wang, Y. Qiao, and X. Tang · 2013
Earlier work this paper cites.
Diverse sequential subset selection for supervised video summarization
B. Gong, W.-L. Chao, K. Grauman, and F. Sha · 2014
Earlier work this paper cites.
A. Graves, G. Wayne, and I. Danihelka · 2014
Earlier work this paper cites.
Action localization with tubelets from motion
M. Jain, J. Van Gemert, H. Jégou, P. Bouthemy, and C. G. Snoek · 2014
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, et al · 2014
Earlier work this paper cites.
Parsing videos of actions with segmental grammars
H. Pirsiavash and D. Ramanan · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Modeling video evolution for action recognition
B. Fernando, E. Gavves, J. M. Oramas, A. Ghodrati, and T. Tuytelaars · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
End-to-end memory networks
S. Sukhbaatar, J. Weston, R. Fergus, et al · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Earlier work this paper cites.
Convolutional two-stream network fusion for video action recognition
C. Feichtenhofer, A. Pinz, and A. Zisserman · 2016
Earlier work this paper cites.
Cross modal distillation for supervision transfer
S. Gupta, J. Hoffman, and J. Malik · 2016
Earlier work this paper cites.
Visually indicated sounds
A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Earlier work this paper cites.
Leaving some stones unturned: dynamic feature prioritization for activity detection in streaming video
Y.-C. Su and K. Grauman · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Cited alongside, same era.
Multi-stream multi-class fusion of deep networks for video classification
Z. Wu, Y.-G. Jiang, X. Wang, H. Ye, and X. Xue · 2016
Cited alongside, same era.
End-to-end learning of action detection from frame glimpses in videos
S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei · 2016
Cited alongside, same era.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
Sst: Single-stream temporal action proposals
S. Buch, V. Escorcia, C. Shen, B. Ghanem, and J. Carlos Niebles · 2017
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Later among the works it cites.
Audio-visual event localization in unconstrained videos
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu · 2018
Later among the works it cites.
A closer look at spatiotemporal convolutions for action recognition
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri · 2018
Later among the works it cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Later among the works it cites.
Compressed video action recognition
C.-Y. Wu, M. Zaheer, H. Hu, R. Manmatha, A. J. Smola, and P. Krähenbühl · 2018
Later among the works it cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
ActionVLAD: Learning spatio-temporal aggregation for action classification
R. Girdhar, D. Ramanan, A. Gupta, J. Sivic, and B. Russell · 2017
Cited alongside, same era.
The kinetics human action video dataset
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al · 2017
Cited alongside, same era.
Unsupervised video summarization with adversarial lstm networks
B. Mahasseni, M. Lam, and S. Todorovic · 2017
Cited alongside, same era.
Learning spatio-temporal representation with pseudo-3d residual networks
Z. Qiu, T. Yao, and T. Mei · 2017
Cited alongside, same era.
Cdc: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos
Z. Shou, J. Chan, A. Zareian, K. Miyazawa, and S.-F. Chang · 2017
Cited alongside, same era.
Retrospective encoders for video summarization
K. Zhang, K. Grauman, and F. Sha · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Later among the works it cites.
Temporal relational reasoning in videos
B. Zhou, A. Andonian, A. Oliva, and A. Torralba · 2018
Later among the works it cites.
Visual to sound: Generating natural sound for videos in the wild
Y. Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg · 2018
Later among the works it cites.
Fine-grained video categorization with redundancy reduction attention
C. Zhu, X. Tan, F. Zhou, X. Liu, K. Yue, E. Ding, and Y. Ma · 2018
Later among the works it cites.
Temporal cycle-consistency learning
D. Dwibedi, Y. Aytar, J. Tompson, P. Sermanet, and A. Zisserman · 2019
Closest in time.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Closest in time.
Self-supervised moving vehicle tracking with stereo sound
C. Gan, H. Zhao, P. Chen, D. Cox, and A. Torralba · 2019
Closest in time.
2.5d visual sound
R. Gao and K. Grauman · 2019
Closest in time.
Co-separating sounds of visual objects
R. Gao and K. Grauman · 2019
Closest in time.
Distinit: Learning video representations without a single labeled video
R. Girdhar, D. Tran, L. Torresani, and D. Ramanan · 2019
Closest in time.
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen · 2019
Closest in time.
Scsampler: Sampling salient clips from video for efficient action recognition
B. Korbar, D. Tran, and L. Torresani · 2019
Closest in time.
Tsm: Temporal shift module for efficient video understanding
J. Lin, C. Gan, and S. Han · 2019
Closest in time.
Learning to localize sound sources in visual scenes: Analysis and applications
A. Senocak, T.-H. Oh, J. Kim, M. Yang, and I. S. Kweon · 2019
Closest in time.
Dmc-net: Generating discriminative motion cues for fast compressed video action recognition
Z. Shou, X. Lin, Y. Kalantidis, L. Sevilla-Lara, M. Rohrbach, S.-F. Chang, and Z. Yan · 2019
Closest in time.
Videobert: A joint model for video and language representation learning
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid · 2019
Closest in time.
What makes training multi-modal networks hard?
W. Wang, D. Tran, and M. Feiszli · 2019
Closest in time.
Long-term feature banks for detailed video understanding
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krahenbuhl, and R. Girshick · 2019
Closest in time.
Multi-agent reinforcement learning based frame sampling for effective untrimmed video recognition
W. Wu, D. He, X. Tan, S. Chen, and S. Wen · 2019
Closest in time.
Adaframe: Adaptive frame selection for fast video recognition
Z. Wu, C. Xiong, C.-Y. Ma, R. Socher, and L. S. Davis · 2019
Closest in time.
Vision-infused deep audio inpainting
H. Zhou, Z. Liu, X. Xu, P. Luo, and X. Wang · 2019
Closest in time.
Cisco visual networking index: Forecast and trends, 2017–2022 white paper
2022
Closest in time.