Fetching the paper…
Reading the bibliography…
This paper deals with the problem of localizing objects in image and video datasets from visual exemplars.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Social interactions: A first-person perspective
Alircza Fathi, Jessica K Hodgins, and James M Rehg · 2012
Earlier work this paper cites.
Ucf101: A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles · 2015
Earlier work this paper cites.
Fast r-cnn
Ross Girshick · 2015
Earlier work this paper cites.
Spatial pyramid pooling in deep convolutional networks for visual recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Fully-convolutional siamese networks for object tracking
Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip HS Torr · 2016
Earlier work this paper cites.
R-fcn: Object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A novel performance evaluation methodology for single-target trackers
Matej Kristan, Jiri Matas, Aleš Leonardis, Tomas Vojir, Roman Pflugfelder, Gustavo Fernandez, Georg Nebehay, Fatih Porikli, and Luka Čehovin · 2016
Earlier work this paper cites.
Detecting engagement in egocentric video
Yu-Chuan Su and Kristen Grauman · 2016
Earlier work this paper cites.
Activities of daily living (adls)
Peter F Edemekong, Deb L Bomgaars, and Shoshana B Levy · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Coco-stuff: Thing and stuff classes in context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari · 2018
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Earlier work this paper cites.
Conditional neural processes
Marta Garnelo, Dan Rosenbaum, Christopher Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo Rezende, and SM Ali Eslami · 2018
Earlier work this paper cites.
Cornernet: Detecting objects as paired keypoints
Hei Law and Jia Deng · 2018
Earlier work this paper cites.
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg · 2018
Earlier work this paper cites.
Trackingnet: A large-scale dataset and benchmark for object tracking in the wild
Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem · 2018
Earlier work this paper cites.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Earlier work this paper cites.
Online multi-object tracking with dual matching attention networks
Ji Zhu, Hua Yang, Nian Liu, Minyoung Kim, Wenjun Zhang, and Ming-Hsuan Yang · 2018
Earlier work this paper cites.
Distractor-aware siamese networks for visual object tracking
Zheng Zhu, Qiang Wang, Bo Li, Wei Wu, Junjie Yan, and Weiming Hu · 2018
Earlier work this paper cites.
Lasot: A high-quality benchmark for large-scale single object tracking
Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling · 2019
Cited alongside, same era.
LVIS: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick · 2019
Cited alongside, same era.
Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Lianghua Huang, Xin Zhao, and Kaiqi Huang · 2019
Cited alongside, same era.
Few-shot object detection via feature reweighting
Bingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu, Jiashi Feng, and Trevor Darrell · 2019
Cited alongside, same era.
Set transformer: A framework for attention-based permutation-invariant neural networks
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh · 2019
Cited alongside, same era.
Gradnet: Gradient-guided network for visual object tracking
Few-shot object detection via association and discrimination
Yuhang Cao, Jiaqi Wang, Ying Jin, Tong Wu, Kai Chen, Ziwei Liu, and Dahua Lin · 2021
Later among the works it cites.
Dual-awareness attention for few-shot object detection
Tung-I Chen, Yueh-Cheng Liu, Hung-Ting Su, Yu-Cheng Chang, Yu-Hsiang Lin, Jia-Fong Yeh, Wen-Chin Chen, and Winston Hsu · 2021
Later among the works it cites.
Query adaptive few-shot object detection with heterogeneous graph convolutional networks
Guangxing Han, Yicheng He, Shiyuan Huang, Jiawei Ma, and Shih-Fu Chang · 2021
Later among the works it cites.
Towards open world object detection
KJ Joseph, Salman Khan, Fahad Shahbaz Khan, and Vineeth N Balasubramanian · 2021
Later among the works it cites.
Baod: Budget-aware object detection
Alejandro Pardo, Mengmeng Xu, Ali Thabet, Pablo Arbelaez, and Bernard Ghanem · 2021
Later among the works it cites.
Defrcn: Decoupled faster r-cnn for few-shot object detection
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peixia Li, Boyu Chen, Wanli Ouyang, Dong Wang, Xiaoyun Yang, and Huchuan Lu · 2019
Cited alongside, same era.
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Cited alongside, same era.
Fcos: Fully convolutional one-stage object detection
Zhi Tian, Chunhua Shen, Hao Chen, and Tong He · 2019
Cited alongside, same era.
Spm-tracker: Series-parallel matching for real-time visual object tracking
Guangting Wang, Chong Luo, Zhiwei Xiong, and Wenjun Zeng · 2019
Cited alongside, same era.
Detectron2
Yuxin Wu, Alexander Kirillov, Francisco Massa, Wan-Yen Lo, and Ross Girshick · 2019
Cited alongside, same era.
Missing labels in object detection
Mengmeng Xu, Yancheng Bai, and Bernard Ghanem · 2019
Cited alongside, same era.
Joint group feature selection and discriminative filter learning for robust visual object tracking
Tianyang Xu, Zhen-Hua Feng, Xiao-Jun Wu, and Josef Kittler · 2019
Cited alongside, same era.
Limeng Qiao, Yuxuan Zhao, Zhiyuan Li, Xi Qiu, Jianan Wu, and Chi Zhang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Learning to track with object permanence
Pavel Tokmakov, Jie Li, Wolfram Burgard, and Adrien Gaidon · 2021
Later among the works it cites.
Meta-detr: Few-shot object detection via unified image-level meta-learning
Gongjie Zhang, Zhipeng Luo, Kaiwen Cui, and Shijian Lu · 2021
Later among the works it cites.
Hallucination improves few-shot object detection
Weilin Zhang and Yu-Xiong Wang · 2021
Later among the works it cites.
Improving multiple object tracking with single object tracking
Linyu Zheng, Ming Tang, Yingying Chen, Guibo Zhu, Jinqiao Wang, and Hanqing Lu · 2021
Later among the works it cites.
Fs-detr: Few-shot detection transformer with prompting and without re-training
Adrian Bulat, Ricardo Guerrero, Brais Martinez, and Georgios Tzimiropoulos · 2022
Closest in time.
Meta faster r-cnn: Towards accurate few-shot object detection with attentive feature alignment
Guangxing Han, Shiyuan Huang, Jiawei Ma, Yicheng He, and Shih-Fu Chang · 2022
Closest in time.
Ego4d: Around the World in 3,000 Hours of Egocentric Video
2022
Closest in time.
Siammask: A framework for fast online object tracking and segmentation
Weiming Hu, Qiang Wang, Li Zhang, Luca Bertinetto, and Philip HS Torr · 2022
Closest in time.
Unified transformer tracker for object tracking
Fan Ma, Mike Zheng Shou, Linchao Zhu, Haoqi Fan, Yilei Xu, Yi Yang, and Zhicheng Yan · 2022
Closest in time.
Localizing objects in 3d from egocentric videos with visual queries, 2022
Jinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao, and Bernard Ghanem · 2022
Closest in time.
Estimating more camera poses for ego-centric videos is essential for vq3d, 2022
Jinjie Mai, Chen Zhao, Abdullah Hamdi, Silvio Giancola, and Bernard Ghanem · 2022
Closest in time.
Hierarchical attention network for few-shot object detection via meta-contrastive learning
Dongwoo Park and Jongmin Lee · 2022
Closest in time.
Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba, Chen Zhao, Silvio Giancola, and Bernard Ghanem · 2022
Closest in time.
Negative frames matter in egocentric visual query 2d localization
Mengmeng Xu, Cheng-Yang Fu, Yanghao Li, Bernard Ghanem, Juan-Manuel Perez-Rua, and Tao Xiang · 2022
Closest in time.
Towards grand unification of object tracking
Bin Yan, Yi Jiang, Peize Sun, Dong Wang, Zehuan Yuan, Ping Luo, and Huchuan Lu · 2022
Closest in time.
Contrastive conditional neural processes
Zesheng Ye and Lina Yao · 2022
Closest in time.
Sylph: A hypernetwork framework for incremental few-shot object detection
Li Yin, JM Perez, and Kevin J Liang · 2022
Closest in time.
Robust multi-object tracking by marginal inference
Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu · 2022
Closest in time.
Tracking objects as pixel-wise distributions
Zelin Zhao, Ze Wu, Yueqing Zhuang, Boxun Li, and Jiaya Jia · 2022
Closest in time.
Global tracking transformers
Xingyi Zhou, Tianwei Yin, Vladlen Koltun, and Philipp Krähenbühl · 2022
Closest in time.
Newsnet: A novel dataset for hierarchical temporal segmentation
Haoqian Wu, Keyu Chen, Haozhe Liu, Mingchen Zhuge, Bing Li, Ruizhi Qiao, Xiujun Shu, Bei Gan, Liangsheng Xu, Bo Ren, Mengmeng Xu, Wentian Zhang, Raghavendra Ramachandra, Chia-Wen Lin, and Bernard Ghanem · 2023
Closest in time.