Fetching the paper…
Reading the bibliography…
Visual Query Localization on long-form egocentric videos requires spatio-temporal search and localization of visually specified objects and is vital to build episodic memory systems.
Organization of memory
Endel Tulving and Wayne Donaldson · 1973
Earlier work this paper cites.
Determining optical flow
Berthold KP Horn and Brian G Schunck · 1981
Earlier work this paper cites.
An iterative image registration technique with an application to stereo vision
Bruce D. Lucas and Takeo Kanade · 1981
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John M. Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Episodic non-markov localization: Reasoning about short-term and long-term features
Joydeep Biswas and Manuela M. Veloso · 2014
Earlier work this paper cites.
Occlusion and motion reasoning for long-term tracking
Yang Hua, Alahari Karteek, and Cordelia Schmid · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun · 2015
Earlier work this paper cites.
End-to-end memory networks
Sainbayar Sukhbaatar, Arthur D. Szlam, Jason Weston, and Rob Fergus · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta · 2015
Earlier work this paper cites.
Fully-convolutional siamese networks for object tracking
Luca Bertinetto, Jack Valmadre, João F. Henriques, Andrea Vedaldi, and Philip H. S. Torr · 2016
Earlier work this paper cites.
Dense Image Correspondences for Computer Vision
Tal Hassner and Ce Liu · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy P. Lillicrap, Koray Kavukcuoglu, and Daan Wierstra · 2016
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Discriminative correlation filter with channel and spatial reliability
Alan Lukezic, Tomas Vojir, Luka Cehovin Zajc, Jiri Matas, and Matej Kristan · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alexander Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
End-to-end representation learning for correlation filter based tracking
Jack Valmadre, Luca Bertinetto, João F. Henriques, Andrea Vedaldi, and Philip H. S. Torr · 2017
Cited alongside, same era.
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
One-shot instance segmentation
Claudio Michaelis, Ivan Ustyuzhaninov, Matthias Bethge, and Alexander S. Ecker · 2018
Cited alongside, same era.
Lasot: A high-quality benchmark for large-scale single object tracking
Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling · 2019
Cited alongside, same era.
Lvis: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollár, and Ross B. Girshick · 2019
Cited alongside, same era.
Adaptive image transformer for one-shot object detection
Ding-Jie Chen, He-Yen Hsieh, and Tyng-Luh Liu · 2021
Later among the works it cites.
Learning target candidate association to keep track of what not to track
Christoph Mayer, Martin Danelljan, Danda Pani Paudel, and Luc Van Gool · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Loftr: Detector-free local feature matching with transformers
Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou · 2021
Later among the works it cites.
Learning spatio-temporal transformer for visual tracking
Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu · 2021
Later among the works it cites.
Mixformer: End-to-end tracking with iterative mixed attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One-shot object detection with co-attention and co-excitation
Ting-I Hsieh, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu · 2019
Cited alongside, same era.
Siamrpn++: Evolution of siamese visual tracking with very deep networks
Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan · 2019
Cited alongside, same era.
Generalized intersection over union: A metric and a loss for bounding box regression
Seyed Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian D. Reid, and Silvio Savarese · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun · 2019
Cited alongside, same era.
Siam r-cnn: Visual tracking by re-detection
Paul Voigtlaender, Jonathon Luiten, Philip H. S. Torr, and B. Leibe · 2019
Cited alongside, same era.
Comparison network for one-shot conditional object detection
Tengfei Zhang, Yue Zhang, Xian Sun, Hao Sun, Menglong Yan, Xue Yang, and Kun Fu · 2019
Cited alongside, same era.
Know your surroundings: Exploiting scene information for object tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte · 2020
Cited alongside, same era.
Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu · 2022
Later among the works it cites.
Semantic-aware fine-grained correspondence
Yingdong Hu, Renhao Wang, Kaifeng Zhang, and Yang Gao · 2022
Later among the works it cites.
Few-view object reconstruction with unknown categories and camera poses
Hanwen Jiang, Zhenyu Jiang, Kristen Grauman, and Yuke Zhu · 2022
Later among the works it cites.
Unified transformer tracker for object tracking
Fan Ma, Mike Zheng Shou, Linchao Zhu, Haoqi Fan, Yilei Xu, Yi Yang, and Zhicheng Yan · 2022
Later among the works it cites.
Trackformer: Multi-object tracking with transformers
Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, and Christoph Feichtenhofer · 2022
Later among the works it cites.
Negative frames matter in egocentric visual query 2d localization
Mengmeng Xu, Cheng-Yang Fu, Yanghao Li, Bernard Ghanem, Juan-Manuel Pérez-Rúa, and Tao Xiang · 2022
Later among the works it cites.
Where is my wallet? modeling object proposal sets for egocentric visual query localization
Mengmeng Xu, Yanghao Li, Cheng-Yang Fu, Bernard Ghanem, Tao Xiang, and Juan-Manuel Pérez-Rúa · 2022
Later among the works it cites.
Tubedetr: Spatio-temporal video grounding with transformers
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Krahenbuhl, and Ishan Misra · 2022
Later among the works it cites.
Global tracking transformers
Xingyi Zhou, Tianwei Yin, Vladlen Koltun, and Philipp Krähenbühl · 2022
Later among the works it cites.
Conceptfusion: Open-set multimodal 3d mapping
Krishna Murthy Jatavallabhula, Ali Kuwajerwala, Qiao Gu, Mohd Omama, Tao Chen, Shuang Li, Ganesh Iyer, Soroush Saryazdi, Nikhil Varma Keetha, Ayush Tewari, Joshua B. Tenenbaum, Celso M. de Melo, M. Krishna, Liam Paull, Florian Shkurti, and Antonio Torralba · 2023
Closest in time.
Lerf: Language embedded radiance fields
Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa, and Matthew Tancik · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth’ee Darcet, Th’eo Moutakanni, Huy Q. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russ Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael G. Rabbat, Vasu Sharma, Gabriel Synnaeve, Huijiao Xu, Hervé Jégou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2023
Closest in time.