Fetching the paper…
Reading the bibliography…
Transformer networks are effective at modeling long-range contextual information and have recently demonstrated exemplary performance in the natural language processing domain.
THUMOS challenge: Action recognition with a large number of classes
Yu-Gang Jiang, Jingen Liu, A Roshan Zamir, George Toderici, Ivan Laptev, Mubarak Shah, and Rahul Sukthankar. 2014 · 2014
Earlier work this paper cites.
The lear submission at thumos 2014
Dan Oneata, Jakob Verbeek, and Cordelia Schmid. 2014 · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos. In NIPS . 568–576
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition . 961–970
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks. In Int. Conf. Comput. Vis
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015 · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4694–4702
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015 · 2015
Earlier work this paper cites.
Daps: Deep action proposals for action understanding. In European Conference on Computer Vision . Springer, 768–784
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Temporal action localization in untrimmed videos via multi-stage cnns. In CVPR
Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016 · 2016
Earlier work this paper cites.
Untrimmed video classification for activity detection: submission to activitynet challenge
Gurkirt Singh and Fabio Cuzzolin. 2016 · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision . Springer, 20–36
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016 · 2016
Earlier work this paper cites.
Cuhk & ethz & siat submission to activitynet challenge 2016
Yuanjun Xiong, Limin Wang, Zhe Wang, Bowen Zhang, Hang Song, Wei Li, Dahua Lin, Yu Qiao, Luc Van Gool, and Xiaoou Tang. 2016 · 2016
Earlier work this paper cites.
Soft-NMS–improving object detection with one line of code. In Proceedings of the IEEE international conference on computer vision . 5561–5569
Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S Davis. 2017 · 2017
Earlier work this paper cites.
End-to-End, Single-Stream Temporal Action Detection in Untrimmed Videos. In BMVC 2017
Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles. 2017 · 2017
Earlier work this paper cites.
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In CVPR
Joao Carreira and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Turn tap: Temporal unit regression network for temporal action proposals. In ICCV
Jiyang Gao, Zhenheng Yang, Chen Sun, Kan Chen, and Ram Nevatia. 2017 · 2017
Earlier work this paper cites.
SCC: Semantic context cascade for efficient action detection. In CVPR
F Caba Heilbron, Wayner Barrios, Victor Escorcia, and Bernard Ghanem. 2017 · 2017
Earlier work this paper cites.
Learning spatio-temporal representation with pseudo-3d residual networks. In Int. Conf. Comput. Vis. 5533–5541
Zhaofan Qiu, Ting Yao, and Tao Mei. 2017 · 2017
Cited alongside, same era.
CDC: convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In CVPR
Zheng Shou, Jonathan Chan, Alireza Zareian, Kazuyuki Miyazawa, and Shih-Fu Chang. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Untrimmednets for weakly supervised action recognition and detection. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition . 4325–4334
Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool. 2017 · 2017
Cited alongside, same era.
R-c3d: Region convolutional 3d network for temporal activity detection. In ICCV
Huijuan Xu, Abir Das, and Kate Saenko. 2017 · 2017
Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid Network
Jialin Gao, Zhixiang Shi, Jiani Li, Guanshuo Wang, Yufeng Yuan, Shiming Ge, and Xi Zhou. 2020 · 2020
Later among the works it cites.
TEA: Temporal Excitation and Aggregation for Action Recognition. In IEEE Conf. Comput. Vis. Pattern Recog. 909–918
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang. 2020 · 2020
Later among the works it cites.
Fast Learning of Temporal Action Proposal via Dense Boundary Generator.. In AAAI . 11499–11506
Chuming Lin, Jian Li, Yabiao Wang, Ying Tai, Donghao Luo, Zhipeng Cui, Chengjie Wang, Jilin Li, Feiyue Huang, and Rongrong Ji. 2020 · 2020
Later among the works it cites.
TEINet: Towards an Efficient Architecture for Video Recognition.. In AAAI . 11669–11676
Zhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Tong Lu. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Temporal action detection with structured segment networks. In ICCV . 2914–2923
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin. 2017 · 2017
Cited alongside, same era.
Rethinking the faster r-cnn architecture for temporal action localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1130–1139
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A Ross, Jia Deng, and Rahul Sukthankar. 2018 · 2018
Cited alongside, same era.
Ctap: Complementary temporal action proposal generation. In Proceedings of the European conference on computer vision (ECCV) . 68–83
Jiyang Gao, Kan Chen, and Ram Nevatia. 2018 · 2018
Cited alongside, same era.
Bsn: Boundary sensitive network for temporal action proposal generation. In Proceedings of the European Conference on Computer Vision (ECCV) . 3–19
Tianwei Lin, Xu Zhao, Haisheng Su, Chongjing Wang, and Ming Yang. 2018 · 2018
Cited alongside, same era.
A Closer Look at Spatiotemporal Convolutions for Action Recognition. In IEEE Conf. Comput. Vis. Pattern Recog
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018 · 2018
Cited alongside, same era.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Eur. Conf. Comput. Vis
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018 · 2018
Cited alongside, same era.
Slowfast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 6202–6211
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019 · 2019
Cited alongside, same era.
Haisheng Su, Weihao Gan, Wei Wu, Junjie Yan, and Yu Qiao. 2020 · 2020
Later among the works it cites.
Dynamic Inference: A New Approach Toward Efficient Video Action Recognition. In Proceedings of CVPR Workshops . 676–677
Wenhao Wu, Dongliang He, Xiao Tan, Shifeng Chen, Yi Yang, and Shilei Wen. 2020 · 2020
Later among the works it cites.
G-TAD: Sub-Graph Localization for Temporal Action Detection. In CVPR . 10156–10165
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem. 2020 · 2020
Later among the works it cites.
Learning texture transformer network for image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 5791–5800
Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. 2020 · 2020
Later among the works it cites.
Is Space-Time Attention All You Need for Video Understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021 · 2021
Closest in time.
Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu. 2021 · 2021
Closest in time.
Manoj Kumar, Dirk Weissenborn, and Nal Kalchbrenner. 2021 · 2021
Closest in time.
TFPose: Direct Human Pose Estimation with Transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Xinlong Wang, and Zhibin Wang. 2021 · 2021
Closest in time.
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. 2021 · 2021
Closest in time.
Temporal Context Aggregation Network for Temporal Action Proposal Refinement
Zhiwu Qing, Haisheng Su, Weihao Gan, Dongliang Wang, Wei Wu, Xiang Wang, Yu Qiao, Junjie Yan, Changxin Gao, and Nong Sang. 2021 · 2021
Closest in time.
Relaxed Transformer Decoders for Direct Action Proposal Generation
Jing Tan, Jiaqi Tang, Limin Wang, and Gangshan Wu. 2021 · 2021
Closest in time.
MVFNet: Multi-View Fusion Network for Efficient Video Recognition. In AAAI
Wenhao Wu, Dongliang He, Tianwei Lin, Fu Li, Chuang Gan, and Errui Ding. 2021 · 2021
Closest in time.
TransCenter: Transformers with Dense Queries for Multiple-Object Tracking
Yihong Xu, Yutong Ban, Guillaume Delorme, Chuang Gan, Daniela Rus, and Xavier Alameda-Pineda. 2021 · 2021
Closest in time.
Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformers
Tianyu Zhu, Markus Hiller, Mahsa Ehsanpour, Rongkai Ma, Tom Drummond, and Hamid Rezatofighi. 2021 · 2021
Closest in time.