Fetching the paper…
Reading the bibliography…
Visual object tracking is a key component to many egocentric vision problems.
A survey of augmented reality
Ronald T Azuma · 1997
Earlier work this paper cites.
A boosted particle filter: Multitarget detection and tracking
Kenji Okuma, Ali Taleghani, De Freitas, J.J. Little, and David Lowe · 2004
Earlier work this paper cites.
Ensemble tracking
S. Avidan · 2005
Earlier work this paper cites.
Social interactions: A first-person perspective
Alircza Fathi, Jessica K. Hodgins, and James M. Rehg · 2012
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Discovering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman · 2012
Earlier work this paper cites.
Detecting activities of daily living in first-person camera views
Hamed Pirsiavash and Deva Ramanan · 2012
Earlier work this paper cites.
Story-driven summarization for egocentric video
Zheng Lu and Kristen Grauman · 2013
Earlier work this paper cites.
Online object tracking: A benchmark
Yi Wu, Jongwoo Lim, and Ming-Hsuan Yang · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Predicting important objects for egocentric video summarization
Yong Jae Lee and Kristen Grauman · 2015
Earlier work this paper cites.
Nus-pro: A new visual tracking challenge
Annan Li, Min Lin, Yi Wu, Ming-Hsuan Yang, and Shuicheng Yan · 2015
Earlier work this paper cites.
Encoding color information for visual tracking: Algorithms and benchmark
Pengpeng Liang, Erik Blasch, and Haibin Ling · 2015
Earlier work this paper cites.
Personal object discovery in first-person videos
Cewu Lu, Renjie Liao, and Jiaya Jia · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Temporal perception and prediction in ego-centric video
Yipin Zhou and Tamara L. Berg · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
A novel performance evaluation methodology for single-target trackers
Matej Kristan, Jiri Matas, Aleš Leonardis, Tomas Vojir, Roman Pflugfelder, Gustavo Fernandez, Georg Nebehay, Fatih Porikli, and Luka Čehovin · 2016
Earlier work this paper cites.
Mot16: A benchmark for multi-object tracking
Anton Milan, Laura Leal-Taixé, Ian Reid, Stefan Roth, and Konrad Schindler · 2016
Earlier work this paper cites.
A benchmark and simulator for uav tracking
Matthias Mueller, Neil Smith, and Bernard Ghanem · 2016
Earlier work this paper cites.
Detecting engagement in egocentric video
Yu-Chuan Su and Kristen Grauman · 2016
Earlier work this paper cites.
Summarization of egocentric videos: A comprehensive survey
Ana Garcia del Molino, Cheston Tan, Joo-Hwee Lim, and Ah-Hwee Tan · 2017
Earlier work this paper cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Seeing invisible poses: Estimating 3d body pose from egocentric video
Hao Jiang and Kristen Grauman · 2017
Earlier work this paper cites.
Need for speed: A benchmark for higher frame rate object tracking
Hamed Kiani Galoogahi, Ashton Fagg, Chen Huang, Deva Ramanan, and Simon Lucey · 2017
Earlier work this paper cites.
The 2017 davis challenge on video object segmentation
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbeláez, Alexander Sorkine-Hornung, and Luc Van Gool · 2017
Earlier work this paper cites.
Ran Tao, Efstratios Gavves, and Arnold WM Smeulders · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2018
Cited alongside, same era.
When will you do what? - anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall · 2018
Cited alongside, same era.
High performance visual tracking with siamese region proposal network
Bo Li, Junjie Yan, Wei Wu, Zheng Zhu, and Xiaolin Hu · 2018
Cited alongside, same era.
Now you see me: evaluating performance in long-term visual tracking
Alan Lukežič, Luka Čehovin Zajc, Tomáš Vojíř, Jiří Matas, and Matej Kristan · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Rolling-unrolling lstms for action anticipation from first-person video
Antonino Furnari and Giovanni Farinella · 2020
Later among the works it cites.
Globaltrack: A simple and strong baseline for long-term tracking
Lianghua Huang, Xin Zhao, and Kaiqi Huang · 2020
Later among the works it cites.
Hota: A higher order metric for evaluating multi-object tracking
Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixe, and Bastian Leibe · 2020
Later among the works it cites.
Understanding human hands in contact at internet scale
Dandan Shan, Jiaqi Geng, Michelle Shu, and David F Fouhey · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long-term visual object tracking benchmark
Abhinav Moudgil and Vineet Gandhi · 2018
Cited alongside, same era.
Trackingnet: A large-scale dataset and benchmark for object tracking in the wild
Matthias Muller, Adel Bibi, Silvio Giancola, Salman Alsubaihi, and Bernard Ghanem · 2018
Cited alongside, same era.
Long-term tracking in the wild: A benchmark
Jack Valmadre, Luca Bertinetto, Joao F Henriques, Ran Tao, Andrea Vedaldi, Arnold WM Smeulders, Philip HS Torr, and Efstratios Gavves · 2018
Cited alongside, same era.
Youtube-vos: Sequence-to-sequence video object segmentation
Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang · 2018
Cited alongside, same era.
Learning discriminative model prediction for tracking
Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte · 2019
Cited alongside, same era.
Atom: Accurate tracking by overlap maximization
Martin Danelljan, Goutam Bhat, Fahad Shahbaz Khan, and Michael Felsberg · 2019
Cited alongside, same era.
Lasot: A high-quality benchmark for large-scale single object tracking
Heng Fan, Liting Lin, Fan Yang, Peng Chu, Ge Deng, Sijia Yu, Hexin Bai, Yong Xu, Chunyuan Liao, and Haibin Ling · 2019
Cited alongside, same era.
Siam r-cnn: Visual tracking by re-detection
Paul Voigtlaender, Jonathon Luiten, Philip HS Torr, and Bastian Leibe · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Transformer tracking
Xin Chen, Bin Yan, Jiawen Zhu, Dong Wang, Xiaoyun Yang, and Huchuan Lu · 2021
Later among the works it cites.
Is first person vision challenging for object tracking?
Matteo Dunnhofer, Antonino Furnari, Giovanni Maria Farinella, and Christian Micheloni · 2021
Later among the works it cites.
Anticipative video transformer
R. Girdhar and K. Grauman · 2021
Later among the works it cites.
Ego-exo: Transferring visual representations from third-person to first-person videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman · 2021
Later among the works it cites.
Learning target candidate association to keep track of what not to track
Christoph Mayer, Martin Danelljan, Danda Pani Paudel, and Luc Van Gool · 2021
Later among the works it cites.
Unidentified video objects: A benchmark for dense, open-world segmentation
Weiyao Wang, Matt Feiszli, Heng Wang, and Du Tran · 2021
Later among the works it cites.
Learning spatio-temporal transformer for visual tracking
Bin Yan, Houwen Peng, Jianlong Fu, Dong Wang, and Huchuan Lu · 2021
Later among the works it cites.
XMem: Long-term video object segmentation with an atkinson-shiffrin memory model
Ho Kei Cheng and Alexander G. Schwing · 2022
Later among the works it cites.
Mixformer: End-to-end tracking with iterative mixed attention
Yutao Cui, Cheng Jiang, Limin Wang, and Gangshan Wu · 2022
Later among the works it cites.
Epic-kitchens visor benchmark: Video segmentations and object relations
Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Amlan Kar, Richard Ely Locke Higgins, Sanja Fidler, David Fouhey, and Dima Damen · 2022
Later among the works it cites.
A survey of embodied ai: From simulators to research tasks
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
1st place solution for youtubevos challenge 2022: Referring video object segmentation, 2022
Zhiwei Hu, Bo Chen, Yuan Gao, Zhilong Ji, and Jinfeng Bai · 2022
Later among the works it cites.
Transforming model prediction for tracking
Christoph Mayer, Martin Danelljan, Goutam Bhat, Matthieu Paul, Danda Pani Paudel, Fisher Yu, and Luc Van Gool · 2022
Later among the works it cites.
Open-world instance segmentation: Exploiting pseudo ground truth from learned pairwise affinity
Weiyao Wang, Matt Feiszli, Heng Wang, Jitendra Malik, and Du Tran · 2022
Later among the works it cites.
Correlation-aware deep tracking
Fei Xie, Chunyu Wang, Guangting Wang, Yue Cao, Wankou Yang, and Wenjun Zeng · 2022
Later among the works it cites.
Mixformer: End-to-end tracking with iterative mixed attention, 2023
Yutao Cui, Cheng Jiang, Gangshan Wu, and Limin Wang · 2023
Closest in time.
Visual object tracking in first person vision
Matteo Dunnhofer, Antonino Furnari, Giovanni Maria Farinella, and Christian Micheloni · 2023
Closest in time.
Tracking everything everywhere all at once
Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely · 2023
Closest in time.