Fetching the paper…
Reading the bibliography…
Video event localization tasks include temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL).
F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity understanding,” in Proceedings of the ieee conference on computer vision and pattern recognition , 2015, pp. 961–970
2015
Earlier work this paper cites.
A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,” Applied Sciences , vol. 6, no. 6, p. 162, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Mesaros, T. Heittola, A. Diment, B. Elizalde, A. Shah, E. Vincent, B. Raj, and T. Virtanen, “Dcase 2017 challenge setup: Tasks, datasets and baseline system,” in DCASE 2017-Workshop on Detection and Classification of Acoustic Scenes and Events , 2017
2017
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6299–6308
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
N. Bodla, B. Singh, R. Chellappa, and L. S. Davis, “Soft-nms–improving object detection with one line of code,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5561–5569
2017
Earlier work this paper cites.
Y. Zhao, Y. Xiong, L. Wang, Z. Wu, X. Tang, and D. Lin, “Temporal action detection with structured segment networks,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 2914–2923
2017
Earlier work this paper cites.
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold et al. , “Cnn architectures for large-scale audio classification,” in 2017 ieee international conference on acoustics, speech and signal processing (icassp) . IEEE, 2017, pp. 131–135
2017
Earlier work this paper cites.
L. Wang, Y. Xiong, D. Lin, and L. Van Gool, “Untrimmednets for weakly supervised action recognition and detection,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 2017, pp. 4325–4334
2017
Earlier work this paper cites.
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks for action recognition in videos,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 11, pp. 2740–2755, 2018
2018
Earlier work this paper cites.
T. Lin, X. Zhao, H. Su, C. Wang, and M. Yang, “Bsn: Boundary sensitive network for temporal action proposal generation,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 3–19
2018
Earlier work this paper cites.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 247–263
2018
Earlier work this paper cites.
Y.-W. Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar, “Rethinking the faster r-cnn architecture for temporal action localization,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 1130–1139
2018
Earlier work this paper cites.
T. Lin, X. Liu, X. Li, E. Ding, and S. Wen, “Bmn: Boundary-matching network for temporal action proposal generation,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 3889–3898
2019
Earlier work this paper cites.
Y. Wu, L. Zhu, Y. Yan, and Y. Yang, “Dual attention matching for audio-visual event localization,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 6292–6300
2019
Earlier work this paper cites.
N. Turpault, R. Serizel, A. Parag Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Workshop on Detection and Classification of Acoustic Scenes and Events , New York City, United States, Oct. 2019. [Online]. Available: https://inria.hal.science/hal-02160855
2019
Earlier work this paper cites.
H. Xu, A. Das, and K. Saenko, “Two-stream region convolutional 3d network for temporal activity detection,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 10, pp. 2319–2332, 2019
2019
Earlier work this paper cites.
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 658–666
2019
Earlier work this paper cites.
R. Zeng, W. Huang, M. Tan, Y. Rong, P. Zhao, J. Huang, and C. Gan, “Graph convolutional networks for temporal action localization,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 7094–7103
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J.-T. Lee, M. Jain, H. Park, and S. Yun, “Cross-attentional audio-visual fusion for weakly-supervised action localization,” in International conference on learning representations , 2020
2020
Earlier work this paper cites.
Y. Bai, Y. Wang, Y. Tong, Y. Yang, Q. Liu, and J. Liu, “Boundary content graph neural network for temporal action proposal generation,” in European Conference on Computer Vision . Springer, 2020, pp. 121–137
2020
Earlier work this paper cites.
M. Xu, C. Zhao, D. S. Rojas, A. Thabet, and B. Ghanem, “G-tad: Sub-graph localization for temporal action detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 156–10 165
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Xu, R. Zeng, Q. Wu, M. Tan, and C. Gan, “Cross-modal relation-aware networks for audio-visual event localization,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 3893–3901
2020
Cited alongside, same era.
H. Xuan, Z. Zhang, S. Chen, J. Yang, and Y. Yan, “Cross-modal attention network for temporal inconsistent audio-visual event localization,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 01, 2020, pp. 279–286
2020
Cited alongside, same era.
Y. Tian, D. Li, and C. Xu, “Unified multisensory perception: Weakly-supervised audio-visual video parsing,” in European Conference on Computer Vision . Springer, 2020, pp. 436–454
2020
Cited alongside, same era.
G. Luo, Y. Zhou, X. Sun, L. Cao, C. Wu, C. Deng, and R. Ji, “Multi-task collaborative network for joint referring expression comprehension and segmentation,” in Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition , 2020, pp. 10 034–10 043
2020
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al. , “Ego4d: Around the world in 3,000 hours of egocentric video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 995–19 012
2022
Later among the works it cites.
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,” Advances in Neural Information Processing Systems , vol. 35, pp. 32 897–32 912, 2022
2022
Later among the works it cites.
B. Mustafa, C. Riquelme, J. Puigcerver, R. Jenatton, and N. Houlsby, “Multimodal contrastive learning with limoe: the language-image mixture of experts,” Advances in Neural Information Processing Systems , vol. 35, pp. 9564–9576, 2022
2022
Later among the works it cites.
Z. Zhu, L. Wang, W. Tang, N. Zheng, and G. Hua, “Contextloc++: A unified context model for temporal action localization,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 8, pp. 9504–9519, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Lu, V. Goswami, M. Rohrbach, D. Parikh, and S. Lee, “12-in-1: Multi-task vision and language representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2020, pp. 10 437–10 446
2020
Cited alongside, same era.
R. Su, D. Xu, L. Sheng, and W. Ouyang, “Pcg-tal: Progressive cross-granularity cooperation for temporal action localization,” IEEE Transactions on Image Processing , vol. 30, pp. 2103–2113, 2020
2020
Cited alongside, same era.
R. Zeng, W. Huang, M. Tan, Y. Rong, P. Zhao, J. Huang, and C. Gan, “Graph convolutional module for temporal action localization in videos,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 10, pp. 6209–6223, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
C. Lin, C. Xu, D. Luo, Y. Wang, Y. Tai, C. Wang, J. Li, F. Huang, and Y. Fu, “Learning salient boundary feature for anchor-free temporal action localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 3320–3329
2021
Cited alongside, same era.
J. Tan, J. Tang, L. Wang, and G. Wu, “Relaxed transformer decoders for direct action proposal generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 13 526–13 535
2021
Cited alongside, same era.
2021
Cited alongside, same era.
B. Duan, H. Tang, W. Wang, Z. Zong, G. Yang, and Y. Yan, “Audio-visual event localization via recursive fusion by joint co-attention,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2021, pp. 4013–4022
2021
Cited alongside, same era.
2023
Later among the works it cites.
Y. Xiao, T. Khandelwal, and R. K. Das, “Fmsg submission for dcase 2023 challenge task 4 on sound event detection with weak labels and synthetic soundscapes,” Proc. DCASE Challenge , 2023
2023
Later among the works it cites.
T. Geng, T. Wang, J. Duan, R. Cong, and F. Zheng, “Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 22 942–22 951
2023
Later among the works it cites.
R. Hebbar, D. Bose, K. Somandepalli, V. Vijai, and S. Narayanan, “A dataset for audio-visual sound event detection in movies,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Shi, Y. Zhong, Q. Cao, L. Ma, J. Li, and D. Tao, “Tridet: Temporal action detection with relative boundary modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 18 857–18 866
2023
Later among the works it cites.
K. K. Rachavarapu et al. , “Boosting positive segments for weakly-supervised audio-visual video parsing,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 10 192–10 202
2023
Later among the works it cites.
W. Su, P. Miao, H. Dou, G. Wang, L. Qiao, Z. Li, and X. Li, “Language adaptive weight generation for multi-task visual grounding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 857–10 866
2023
Later among the works it cites.
K. Q. Lin, P. Zhang, J. Chen, S. Pramanick, D. Gao, A. J. Wang, R. Yan, and M. Z. Shou, “Univtg: Towards unified video-language temporal grounding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2794–2804
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 15 180–15 190
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
H.-J. Kim, J.-H. Hong, H. Kong, and S.-W. Lee, “Te-tad: Towards full end-to-end temporal action detection via time-aligned coordinate expression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 18 837–18 846
2024
Closest in time.
M. Yang, H. Gao, P. Guo, and L. Wang, “Adapting short-term transformers for action detection in untrimmed videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 18 570–18 579
2024
Closest in time.
L. G. Foo, T. Li, H. Rahmani, and J. Liu, “Action detection via an image diffusion process,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 18 351–18 361
2024
Closest in time.
2024
Closest in time.
W. Hou, G. Li, Y. Tian, and D. Hu, “Toward long form audio-visual video understanding,” ACM Transactions on Multimedia Computing, Communications and Applications , vol. 20, no. 9, pp. 1–26, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
S. Chowdhury, S. Nag, S. Dasgupta, J. Chen, M. Elhoseiny, R. Gao, and D. Manocha, “Meerkat: Audio-visual large language model for grounding in space and time,” in European Conference on Computer Vision . Springer, 2024, pp. 52–70
2024
Closest in time.