Fetching the paper…
Reading the bibliography…
In this report, we present our approach for the Natural Language Query track and Goal Step track of the Ego4D Episodic Memory Benchmark at CVPR 2024.
Lvis: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick · 2019
Earlier work this paper cites.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Multimodal dialog system: Relational graph-based context-aware question understanding
Haoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei, Yinglong Wang, and Liqiang Nie · 2021
Earlier work this paper cites.
Internvideo-ego4d: A pack of champion solutions to ego4d challenges
Guo Chen, Sen Xing, Zhe Chen, Yi Wang, Kunchang Li, Yizhuo Li, Yi Liu, Jiahao Wang, Yin-Dong Zheng, Bingkun Huang, et al · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Earlier work this paper cites.
Bi-directional heterogeneous graph hashing towards efficient outfit recommendation
Weili Guan, Xuemeng Song, Haoyu Zhang, Meng Liu, Chung-Hsing Yeh, and Xiaojun Chang · 2022
Cited alongside, same era.
Egocentric video-language pretraining@ ego4d challenge 2022
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Cited alongside, same era.
Reler@ zju-alibaba submission to the ego4d natural language queries challenge 2022
Naiyuan Liu, Xiaohan Wang, Xiaobo Li, Yi Yang, and Yueting Zhuang · 2022
Cited alongside, same era.
A simple transformer-based model for ego4d natural language queries challenge
Sicheng Mo, Fangzhou Mu, and Yin Li · 2022
Cited alongside, same era.
ActionFormer: Localizing moments of actions with transformers
Naq: Leveraging narrations as queries to supervise episodic memory
Santhosh Kumar Ramakrishnan, Ziad Al-Halah, and Kristen Grauman · 2023
Later among the works it cites.
Action sensitivity learning for the ego4d episodic memory challenge 2023
Jiayi Shao, Xiaohan Wang, Ruijie Quan, and Yi Yang · 2023
Later among the works it cites.
Detrs with collaborative hybrid assignments training
Zhuofan Zong, Guanglu Song, and Yu Liu · 2023
Later among the works it cites.
Ego4d goal-step: Toward hierarchical understanding of procedural activities
Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani · 2024
Closest in time.
Multi-factor adaptive vision selection for egocentric video question answering
Haoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song, Yaowei Wang, and Liqiang Nie · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen-Lin Zhang, Jianxin Wu, and Yin Li · 2022
Cited alongside, same era.
Groundnlq@ ego4d natural language queries challenge 2023
Zhijian Hou, Lei Ji, Difei Gao, Wanjun Zhong, Kun Yan, Chao Li, Wing-Kwong Chan, Chong-Wah Ngo, Nan Duan, and Mike Zheng Shou · 2023
Cited alongside, same era.
CONE: An efficient coarse-to-fiNE alignment framework for long video temporal grounding
Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, Wk Chan, Chong-Wah Ngo, Mike Zheng Shou, and Nan Duan
Cited in the paper.
Attribute-guided collaborative learning for partial person re-identification
Haoyu Zhang, Meng Liu, Yuhong Li, Ming Yan, Zan Gao, Xiaojun Chang, and Liqiang Nie
Cited in the paper.
Uncovering hidden connections: Iterative tracking and reasoning for video-grounded dialog
Haoyu Zhang, Meng Liu, Yaowei Wang, Da Cao, Weili Guan, and Liqiang Nie
Cited in the paper.