Fetching the paper…
Reading the bibliography…
Self-attention based Transformer models have demonstrated impressive results for image classification and object detection, and more recently for video understanding.
Caba Heilbron, F., Escorcia, V., Ghanem, B., Carlos Niebles, J.: ActivityNet: A large-scale video benchmark for human activity understanding. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 961–970 (2015)
2015
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: Int. Conf. Learn. Represent. pp. 1–11 (2015)
2015
Earlier work this paper cites.
Caba Heilbron, F., Carlos Niebles, J., Ghanem, B.: Fast temporal activity proposals for efficient detection of human actions in untrimmed videos. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1914–1923 (2016)
2016
Earlier work this paper cites.
Escorcia, V., Heilbron, F.C., Niebles, J.C., Ghanem, B.: DAPs: Deep action proposals for action understanding. In: Eur. Conf. Comput. Vis. LNCS, vol. 9907, pp. 768–784 (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 770–778 (2016)
2016
Earlier work this paper cites.
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C.: SSD: Single shot multibox detector. In: Eur. Conf. Comput. Vis. pp. 21–37 (2016)
2016
Earlier work this paper cites.
Shou, Z., Wang, D., Chang, S.F.: Temporal action localization in untrimmed videos via multi-stage CNNs. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1049–1058 (2016)
2016
Earlier work this paper cites.
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Van Gool, L.: Temporal segment networks: Towards good practices for deep action recognition. In: Eur. Conf. Comput. Vis. LNCS, vol. 11205, pp. 20–36 (2016)
2016
Earlier work this paper cites.
Bodla, N., Singh, B., Chellappa, R., Davis, L.S.: Soft-NMS–improving object detection with one line of code. In: Int. Conf. Comput. Vis. pp. 5561–5569 (2017)
2017
Earlier work this paper cites.
Buch, S., Escorcia, V., Ghanem, B., Niebles Carlos, J.: End-to-end, single-stream temporal action detection in untrimmed videos. In: Brit. Mach. Vis. Conf. pp. 93.1–93.12 (2017)
2017
Earlier work this paper cites.
Buch, S., Escorcia, V., Shen, C., Ghanem, B., Carlos Niebles, J.: SST: Single-stream temporal action proposals. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2911–2920 (2017)
2017
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the Kinetics dataset. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4724–4733 (2017)
2017
Earlier work this paper cites.
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J.E., Weinberger, K.Q.: Snapshot ensembles: Train 1, get m for free. In: Int. Conf. Learn. Represent. (2017)
2017
Earlier work this paper cites.
Idrees, H., Zamir, A.R., Jiang, Y.G., Gorban, A., Laptev, I., Sukthankar, R., Shah, M.: The THUMOS challenge on action recognition for videos “in the wild”. Comput. Vis. and Image Under. 155
2017
Earlier work this paper cites.
Lin, T., Zhao, X., Shou, Z.: Single shot temporal action detection. In: ACM Int. Conf. Multimedia. pp. 988–996 (2017)
2017
Earlier work this paper cites.
Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 2117–2125 (2017)
2017
Earlier work this paper cites.
Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: Int. Conf. Comput. Vis. pp. 2980–2988 (2017)
2017
Earlier work this paper cites.
Qiu, Z., Yao, T., Mei, T.: Learning spatio-temporal representation with pseudo-3d residual networks. In: Int. Conf. Comput. Vis. pp. 5533–5541 (2017)
2017
Earlier work this paper cites.
Shou, Z., Chan, J., Zareian, A., Miyazawa, K., Chang, S.F.: CDC: Convolutional-de-convolutional networks for precise temporal action localization in untrimmed videos. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 5734–5743 (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Adv. Neural Inform. Process. Syst. pp. 5998–6008 (2017)
2017
Earlier work this paper cites.
Zhao, Y., Xiong, Y., Wang, L., Wu, Z., Tang, X., Lin, D.: Temporal action detection with structured segment networks. In: Int. Conf. Comput. Vis. pp. 2914–2923 (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Alwassel, H., Caba Heilbron, F., Escorcia, V., Ghanem, B.: Diagnosing error in temporal action detectors. In: Eur. Conf. Comput. Vis. LNCS, vol. 11207, pp. 256–272 (2018)
2018
Earlier work this paper cites.
Chao, Y.W., Vijayanarasimhan, S., Seybold, B., Ross, D.A., Deng, J., Sukthankar, R.: Rethinking the Faster-RCNN architecture for temporal action localization. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1130–1139 (2018)
2018
Earlier work this paper cites.
Lin, T., Zhao, X., Su, H., Wang, C., Yang, M.: BSN: Boundary sensitive network for temporal action proposal generation. In: Eur. Conf. Comput. Vis. LNCS, vol. 11208, pp. 3–19 (2018)
2018
Earlier work this paper cites.
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: A closer look at spatiotemporal convolutions for action recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6450–6459 (2018)
2018
Earlier work this paper cites.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: North American Asso. Comput. Lin. pp. 4171–4186 (2019)
2019
Earlier work this paper cites.
Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., Tian, Q.: CenterNet: Keypoint triplets for object detection. In: Int. Conf. Comput. Vis. pp. 6569–6578 (2019)
2019
Earlier work this paper cites.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: SlowFast networks for video recognition. In: Int. Conf. Comput. Vis. pp. 6202–6211 (2019)
2019
Earlier work this paper cites.
Girdhar, R., Carreira, J., Doersch, C., Zisserman, A.: Video action transformer network. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 244–253 (2019)
2019
Earlier work this paper cites.
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.X., Yan, X.: Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In: Adv. Neural Inform. Process. Syst. vol. 32 (2019)
2019
Cited alongside, same era.
Lin, T., Liu, X., Li, X., Ding, E., Wen, S.: BMN: Boundary-matching network for temporal action proposal generation. In: Int. Conf. Comput. Vis. pp. 3889–3898 (2019)
2019
Cited alongside, same era.
Liu, D., Jiang, T., Wang, Y.: Completeness modeling and context separation for weakly supervised temporal action localization. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1298–1307 (2019)
2019
Cited alongside, same era.
Liu, Y., Ma, L., Zhang, Y., Liu, W., Chang, S.F.: Multi-granularity generator for temporal action proposal. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3604–3613 (2019)
2019
Cited alongside, same era.
Dai, X., Chen, Y., Yang, J., Zhang, P., Yuan, L., Zhang, L.: Dynamic DETR: End-to-end object detection with dynamic attention. In: Int. Conf. Comput. Vis. pp. 2988–2997 (2021)
2021
Later among the works it cites.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: Int. Conf. Learn. Represent. (2021)
2021
Later among the works it cites.
Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., Feichtenhofer, C.: Multiscale vision transformers. In: Int. Conf. Comput. Vis. (2021)
2021
Later among the works it cites.
Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q.V., Sung, Y., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. In: Int. Conf. Mach. Learn. (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long, F., Yao, T., Qiu, Z., Tian, X., Luo, J., Mei, T.: Gaussian temporal awareness networks for action localization. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 344–353 (2019)
2019
Cited alongside, same era.
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al.: PyTorch: An imperative style, high-performance deep learning library. In: Adv. Neural Inform. Process. Syst. vol. 32 (2019)
2019
Cited alongside, same era.
Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., Savarese, S.: Generalized intersection over union: A metric and a loss for bounding box regression. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 658–666 (2019)
2019
Cited alongside, same era.
Tian, Z., Shen, C., Chen, H., He, T.: FCOS: Fully convolutional one-stage object detection. In: Int. Conf. Comput. Vis. pp. 9627–9636 (2019)
2019
Cited alongside, same era.
Zeng, R., Huang, W., Tan, M., Rong, Y., Zhao, P., Huang, J., Gan, C.: Graph convolutional networks for temporal action localization. In: Int. Conf. Comput. Vis. pp. 7094–7103 (2019)
2019
Cited alongside, same era.
Bai, Y., Wang, Y., Tong, Y., Yang, Y., Liu, Q., Liu, J.: Boundary content graph neural network for temporal action proposal generation. In: Eur. Conf. Comput. Vis. LNCS, vol. 12373, pp. 121–137 (2020)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: Eur. Conf. Comput. Vis. LNCS, vol. 12346, pp. 213–229 (2020)
2020
Cited alongside, same era.
Lin, C., Xu, C., Luo, D., Wang, Y., Tai, Y., Wang, C., Li, J., Huang, F., Fu, Y.: Learning salient boundary feature for anchor-free temporal action localization. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3320–3329 (2021)
2021
Later among the works it cites.
Liu, X., Hu, Y., Bai, S., Ding, F., Bai, X., Torr, P.H.: Multi-shot temporal event localization: a benchmark. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 12596–12606 (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Int. Conf. Comput. Vis. (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Qing, Z., Su, H., Gan, W., Wang, D., Wu, W., Wang, X., Qiao, Y., Yan, J., Gao, C., Sang, N.: Temporal context aggregation network for temporal action proposal refinement. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 485–494 (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. (2021)
2021
Later among the works it cites.
Sridhar, D., Quader, N., Muralidharan, S., Li, Y., Dai, P., Lu, J.: Class semantics-based attention for action detection. In: Int. Conf. Comput. Vis. pp. 13739–13748 (2021)
2021
Later among the works it cites.
Tan, J., Tang, J., Wang, L., Wu, G.: Relaxed transformer decoders for direct action proposal generation. In: Int. Conf. Comput. Vis. pp. 13526–13535 (2021)
2021
Later among the works it cites.
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H.: Training data-efficient image transformers & distillation through attention. In: Int. Conf. Mach. Learn. pp. 10347–10357 (2021)
2021
Later among the works it cites.
Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., Jégou, H.: Going deeper with image transformers. In: Int. Conf. Comput. Vis. (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Wang, T., Yuan, L., Chen, Y., Feng, J., Yan, S.: PnP-DETR: Towards efficient visual analysis with transformers. In: Int. Conf. Comput. Vis. pp. 4661–4670 (2021)
2021
Later among the works it cites.
Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: Int. Conf. Comput. Vis. (2021)
2021
Later among the works it cites.
Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., Xia, H.: End-to-end video instance segmentation with Transformers. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 8741–8750 (2021)
2021
Later among the works it cites.
Xiao, T., Singh, M., Mintun, E., Darrell, T., Dollár, P., Girshick, R.: Early convolutions help transformers see better. In: Adv. Neural Inform. Process. Syst. (2021)
2021
Later among the works it cites.
Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., Luo, P.: Segformer: Simple and efficient design for semantic segmentation with transformers. In: Adv. Neural Inform. Process. Syst. (2021)
2021
Later among the works it cites.
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., Singh, V.: Nyströmformer: A nyström-based algorithm for approximating self-attention. In: AAAI. vol. 35, pp. 14138–14148 (2021)
2021
Later among the works it cites.
Xu, M., Pérez-Rúa, J.M., Escorcia, V., Martinez, B., Zhu, X., Zhang, L., Ghanem, B., Xiang, T.: Boundary-sensitive pre-training for temporal localization in videos. In: Int. Conf. Comput. Vis. pp. 7220–7230 (2021)
2021
Later among the works it cites.
Yang, J., Li, C., Zhang, P., Dai, X., Xiao, B., Yuan, L., Gao, J.: Focal self-attention for local-global interactions in vision transformers. In: Adv. Neural Inform. Process. Syst. (2021)
2021
Later among the works it cites.
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z., Tay, F.E., Feng, J., Yan, S.: Tokens-to-token ViT: Training vision transformers from scratch on ImageNet. In: Int. Conf. Comput. Vis. (2021)
2021
Later among the works it cites.
Zhao, C., Thabet, A.K., Ghanem, B.: Video self-stitching graph network for temporal action localization. In: Int. Conf. Comput. Vis. pp. 13658–13667 (2021)
2021
Later among the works it cites.
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable DETR: Deformable transformers for end-to-end object detection. In: Int. Conf. Learn. Represent. pp. 1–11 (2021)
2021
Later among the works it cites.
Zhu, Z., Tang, W., Wang, L., Zheng, N., Hua, G.: Enriching local and global contexts for temporal action localization. In: Int. Conf. Comput. Vis. pp. 13516–13525 (2021)
2021
Later among the works it cites.
Yang, Z., Qin, J., Huang, D.: Acgnet: Action complement graph network for weakly-supervised temporal action localization. In: AAAI. vol. 36-3, pp. 3090–3098 (2022)
2022
Closest in time.