Fetching the paper…
Reading the bibliography…
Current methods for Video Moment Retrieval (VMR) struggle to align complex situations involving specific environmental details, character descriptions, and action narratives.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14 . Springer, 2016, pp. 510–526
2016
Earlier work this paper cites.
L. Anne Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5803–5812
2017
Earlier work this paper cites.
J. Gao, C. Sun, Z. Yang, and R. Nevatia, “Tall: Temporal activity localization via language query,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5267–5275
2017
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6077–6086
2018
Earlier work this paper cites.
Y. Yuan, L. Ma, J. Wang, W. Liu, and W. Zhu, “Semantic conditioned dynamic modulation for temporal sentence grounding in videos,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
D. Zhang, X. Dai, X. Wang, Y.-F. Wang, and L. S. Davis, “Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 1247–1257
2019
Earlier work this paper cites.
Z. Zhang, Z. Lin, Z. Zhao, and Z. Xiao, “Cross-modal interaction networks for query-based moment retrieval in videos,” in Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval , 2019, pp. 655–664
2019
Earlier work this paper cites.
R. Ge, J. Gao, K. Chen, and R. Nevatia, “Mac: Mining activity concepts for language-based temporal localization,” in 2019 IEEE winter conference on applications of computer vision (WACV) . IEEE, 2019, pp. 245–253
2019
Earlier work this paper cites.
S. Chen and Y.-G. Jiang, “Semantic proposal for activity localization in videos via sentence query,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 8199–8206
2019
Earlier work this paper cites.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 6202–6211
2019
Earlier work this paper cites.
H. Rezatofighi, N. Tsoi, J. Gwak, A. Sadeghian, I. Reid, and S. Savarese, “Generalized intersection over union: A metric and a loss for bounding box regression,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 658–666
2019
Earlier work this paper cites.
S. Chen, W. Jiang, W. Liu, and Y.-G. Jiang, “Learning modality interaction for temporal sentence localization and event captioning in videos,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 . Springer, 2020, pp. 333–351
2020
Earlier work this paper cites.
X. Qu, P. Tang, Z. Zou, Y. Cheng, J. Dong, P. Zhou, and Z. Xu, “Fine-grained iterative attention network for temporal language localization in videos,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 4280–4288
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
D. Liu, X. Qu, X.-Y. Liu, J. Dong, P. Zhou, and Z. Xu, “Jointly cross-and self-modal graph attention network for query-based moment localization,” in Proceedings of the 28th ACM International Conference on Multimedia , 2020, pp. 4070–4078
2020
Earlier work this paper cites.
J. Lei, L. Yu, T. L. Berg, and M. Bansal, “Tvr: A large-scale dataset for video-subtitle moment retrieval,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16 . Springer, 2020, pp. 447–463
2020
Earlier work this paper cites.
S. Zhang, H. Peng, J. Fu, and J. Luo, “Learning 2d temporal adjacent networks for moment localization with natural language,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 12 870–12 877
2020
Cited alongside, same era.
M. Zhang, Y. Yang, X. Chen, Y. Ji, X. Xu, J. Li, and H. T. Shen, “Multi-stage aggregated transformer network for temporal language localization in videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 669–12 678
2021
Cited alongside, same era.
Y.-W. Chen, Y.-H. Tsai, and M.-H. Yang, “End-to-end multi-modal video temporal grounding,” Advances in Neural Information Processing Systems , vol. 34, pp. 28 442–28 453, 2021
2021
Cited alongside, same era.
H. Wang, Z.-J. Zha, L. Li, D. Liu, and J. Luo, “Structured multi-level interaction network for video moment localization via language query,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7026–7035
O. OpenAI, “Gpt-4 technical report,” Mar 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
J. Lei, T. L. Berg, and M. Bansal, “Detecting moments and highlights in videos via natural language queries,” Advances in Neural Information Processing Systems , vol. 34, pp. 11 846–11 858, 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
J. Gao and C. Xu, “Fast video moment retrieval,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1523–1532
2021
Cited alongside, same era.
D. Liu, X. Qu, J. Dong, P. Zhou, Y. Cheng, W. Wei, Z. Xu, and Y. Xie, “Context-aware biaffine localizing network for temporal sentence grounding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 11 235–11 244
2021
Cited alongside, same era.
Y. Zeng, D. Cao, X. Wei, M. Liu, Z. Zhao, and Z. Qin, “Multi-modal relational graph for cross-modal video moment retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 2215–2224
2021
Cited alongside, same era.
Y. Liu, S. Li, Y. Wu, C.-W. Chen, Y. Shan, and X. Qie, “Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 3042–3051
2022
Cited alongside, same era.
J. Hao, H. Sun, P. Ren, J. Wang, Q. Qi, and J. Liao, “Query-aware video encoder for video moment retrieval,” Neurocomputing , vol. 483, pp. 72–86, 2022
2022
Cited alongside, same era.
C. Ju, T. Han, K. Zheng, Y. Zhang, and W. Xie, “Prompting visual-language models for efficient video understanding,” in European Conference on Computer Vision . Springer, 2022, pp. 105–124
2022
Cited alongside, same era.
2023
Later among the works it cites.
M. F. Naeem, M. G. Z. A. Khan, Y. Xian, M. Z. Afzal, D. Stricker, L. Van Gool, and F. Tombari, “I2mvformer: Large language model generated multi-view document supervision for zero-shot image classification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 15 169–15 179
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Li, M. Cao, X. Cheng, Y. Li, Z. Zhu, and Y. Zou, “G2l: Semantically aligned and uniform video grounding via geodesic and game theory,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 12 032–12 042
2023
Later among the works it cites.
H. Li, X. Shu, S. He, R. Qiao, W. Wen, T. Guo, B. Gan, and X. Sun, “D3g: Exploring gaussian prior for temporal sentence grounding with glance annotation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 13 734–13 746
2023
Later among the works it cites.
S. Yan, X. Xiong, A. Nagrani, A. Arnab, Z. Wang, W. Ge, D. Ross, and C. Schmid, “Unloc: A unified framework for video localization tasks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 13 623–13 633
2023
Later among the works it cites.
K. Q. Lin, P. Zhang, J. Chen, S. Pramanick, D. Gao, A. J. Wang, R. Yan, and M. Z. Shou, “Univtg: Towards unified video-language temporal grounding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2794–2804
2023
Later among the works it cites.
S. Yu, J. Cho, P. Yadav, and M. Bansal, “Self-chained image-language model for video localization and question answering,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.