Fetching the paper…
Reading the bibliography…
This report presents ReLER submission to two tracks in the Ego4D Episodic Memory Benchmark in CVPR 2023, including Natural Language Queries and Moment Queries.
Soft-nms–improving object detection with one line of code
Navaneeth Bodla, Bharat Singh, Rama Chellappa, and Larry S Davis · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He · 2019
Earlier work this paper cites.
Bmn: Boundary-matching network for temporal action proposal generation
Tianwei Lin, Xiao Liu, Xin Li, Errui Ding, and Shilei Wen · 2019
Earlier work this paper cites.
G-tad: Sub-graph localization for temporal action detection
Mengmeng Xu, Chen Zhao, David S Rojas, Ali Thabet, and Bernard Ghanem · 2020
Earlier work this paper cites.
Span-based localizing network for natural language video localization
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou · 2020
Earlier work this paper cites.
Learning 2d temporal adjacent networks for moment localization with natural language
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo · 2020
Earlier work this paper cites.
Distance-iou loss: Faster and better learning for bounding box regression
Zhaohui Zheng, Ping Wang, Wei Liu, Jinze Li, Rongguang Ye, and Dongwei Ren · 2020
Earlier work this paper cites.
Detecting moments and highlights in videos via natural language queries
Jie Lei, Tamara L Berg, and Mohit Bansal · 2021
Cited alongside, same era.
Learning salient boundary feature for anchor-free temporal action localization
Chuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Interactive prototype learning for egocentric action recognition
Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang · 2021
Cited alongside, same era.
Internvideo-ego4d: A pack of champion solutions to ego4d challenges
Guo Chen, Sen Xing, Zhe Chen, Yi Wang, Kunchang Li, Yizhuo Li, Yi Liu, Jiahao Wang, Yin-Dong Zheng, Bingkun Huang, et al · 2022
Cited alongside, same era.
Reler@ zju-alibaba submission to the ego4d natural language queries challenge 2022
Naiyuan Liu, Xiaohan Wang, Xiaobo Li, Yi Yang, and Yueting Zhuang · 2022
Later among the works it cites.
A simple transformer-based model for ego4d natural language queries challenge
Sicheng Mo, Fangzhou Mu, and Yin Li · 2022
Later among the works it cites.
Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision
Xiaohan Wang, Linchao Zhu, Zhedong Zheng, Mingliang Xu, and Yi Yang · 2022
Later among the works it cites.
Actionformer: Localizing moments of actions with transformers
Chen-Lin Zhang, Jianxin Wu, and Yin Li · 2022
Later among the works it cites.
Centerclip: Token clustering for efficient text-video retrieval
Shuai Zhao, Linchao Zhu, Xiaohan Wang, and Yi Yang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Omnivore: A single model for many visual modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Cited alongside, same era.
Egocentric video-language pretraining
Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al · 2022
Cited alongside, same era.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang
Cited in the paper.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai
Cited in the paper.
Naq: Leveraging narrations as queries to supervise episodic memory
Santhosh K. Ramakrishnan, Ziad Al-Halah, and Kristen Grauman · 2023
Closest in time.
Action sensitivity learning for temporal action localization
Jiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng, Jiang Yang, and Yi Yang · 2023
Closest in time.