Fetching the paper…
Reading the bibliography…
Referring understanding is a fundamental task that bridges natural language and visual content by localizing objects described in free-form expressions.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Modeling context between objects for referring expression understanding
Varun K Nagaraja, Vlad I Morariu, and Larry S Davis · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Mot16: A benchmark for multi-object tracking
Anton Milan, Laura Leal-Taixé, Ian Reid, Stefan Roth, and Konrad Schindler · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Tracking by natural language specification
Zhenyang Li, Ran Tao, Efstratios Gavves, Cees GM Snoek, and Arnold WM Smeulders · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Simple online and realtime tracking with a deep association metric
Nicolai Wojke, Alex Bewley, and Dietrich Paulus · 2017
Earlier work this paper cites.
Object referring in videos with language and human gaze
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool · 2018
Earlier work this paper cites.
Actor and action video segmentation from a sentence
Kirill Gavrilyuk, Amir Ghodrati, Zhenyang Li, and Cees GM Snoek · 2018
Earlier work this paper cites.
Talk2car: Taking control of your self-driving car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens · 2019
Earlier work this paper cites.
Video object segmentation with language referring expressions
Anna Khoreva, Anna Rohrbach, and Bernt Schiele · 2019
Earlier work this paper cites.
Weakly-supervised spatio-temporally grounding natural sentence in video
Zhenfang Chen, Lin Ma, Wenhan Luo, and Kwan-Yee K Wong · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Relationship-embedded representation learning for grounding referring expressions
Sibei Yang, Guanbin Li, and Yizhou Yu · 2020
Earlier work this paper cites.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han · 2020
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Earlier work this paper cites.
Cops-ref: A new dataset and task on compositional referring expression comprehension
Zhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K Wong, and Qi Wu · 2020
Earlier work this paper cites.
Tao: A large-scale benchmark for tracking any object
Achal Dave, Tarasha Khurana, Pavel Tokmakov, Cordelia Schmid, and Deva Ramanan · 2020
Earlier work this paper cites.
Mot20: A benchmark for multi object tracking in crowded scenes
Patrick Dendorfer, Hamid Rezatofighi, Anton Milan, Javen Shi, Daniel Cremers, Ian Reid, Stefan Roth, Konrad Schindler, and Laura Leal-Taixé · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Transtrack: Multiple object tracking with transformer
Peize Sun, Jinkun Cao, Yi Jiang, Rufeng Zhang, Enze Xie, Zehuan Yuan, Changhu Wang, and Ping Luo · 2020
Cited alongside, same era.
Hota: A higher order metric for evaluating multi-object tracking
Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taixé, and Bastian Leibe · 2020
Cited alongside, same era.
Bevdet: High-performance multi-camera 3d object detection in bird-eye-view
Junjie Huang, Guan Huang, Zheng Zhu, Yun Ye, and Dalong Du · 2021
Cited alongside, same era.
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Cited alongside, same era.
Motiontrack: Learning robust short-term and long-term motions for multi-object tracking
Zheng Qin, Sanping Zhou, Le Wang, Jinghai Duan, Gang Hua, and Wei Tang · 2023
Later among the works it cites.
Tcovis: Temporally consistent online video instance segmentation
Junlong Li, Bingyao Yu, Yongming Rao, Jie Zhou, and Jiwen Lu · 2023
Later among the works it cites.
Memotr: Long-term memory-augmented transformer for multi-object tracking
Ruopeng Gao and Limin Wang · 2023
Later among the works it cites.
https://chat.openai.com , 2023
OpenAI · 2023
Later among the works it cites.
Motrv2: Bootstrapping end-to-end multi-object tracking by pretrained object detectors
Yuang Zhang, Tiancai Wang, and Xiangyu Zhang · 2023
Later among the works it cites.
ikun: Speak to trackers without retraining
Yunhao Du, Cheng Lei, Zhicheng Zhao, and Fei Su · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fairmot: On the fairness of detection and re-identification in multiple object tracking
Yifu Zhang, Chunyu Wang, Xinggang Wang, Wenjun Zeng, and Wenyu Liu · 2021
Cited alongside, same era.
Motr: End-to-end multiple-object tracking with transformer
Fangao Zeng, Bin Dong, Yuang Zhang, Tiancai Wang, Xiangyu Zhang, and Yichen Wei · 2022
Cited alongside, same era.
Dancetrack: Multi-object tracking in uniform appearance and diverse motion
Peize Sun, Jinkun Cao, Yi Jiang, Zehuan Yuan, Song Bai, Kris Kitani, and Ping Luo · 2022
Cited alongside, same era.
Transvod: End-to-end video object detection with spatial-temporal transformers
Qianyu Zhou, Xiangtai Li, Lu He, Yibo Yang, Guangliang Cheng, Yunhai Tong, Lizhuang Ma, and Dacheng Tao · 2022
Cited alongside, same era.
Memot: Multi-object tracking with memory
Jiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong, Wei Xia, Zhuowen Tu, and Stefano Soatto · 2022
Cited alongside, same era.
Inspro: Propagating instance query and proposal for online video instance segmentation
Fei He, Haoyang Zhang, Naiyu Gao, Jian Jia, Yanhu Shan, Xin Zhao, and Kaiqi Huang · 2022
Cited alongside, same era.
Temporally efficient vision transformer for video instance segmentation
Shusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang, Jiemin Fang, Wenyu Liu, Xun Zhao, and Ying Shan · 2022
Cited alongside, same era.
Closest in time.
Echotrack: Auditory referring multi-object tracking for autonomous driving
Jiacheng Lin, Jiajun Chen, Kunyu Peng, Xuan He, Zhiyong Li, Rainer Stiefelhagen, and Kailun Yang · 2024
Closest in time.
Visual-linguistic representation learning with deep cross-modality fusion for referring multi-object tracking
Wenyan He, Yajun Jian, Yang Lu, and Hanzi Wang · 2024
Closest in time.
Mls-track: Multilevel semantic interaction in rmot
Zeliang Ma, Song Yang, Zhe Cui, Zhicheng Zhao, Fei Su, Delong Liu, and Jingyu Wang · 2024
Closest in time.
Wenjun Huang, Yang Ni, Hanning Chen, Yirui He, Ian Bryant, Yezi Liu, and Mohsen Imani · 2024
Closest in time.
Bevformer: learning bird’s-eye-view representation from lidar-camera via spatiotemporal transformers
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai · 2024
Closest in time.
Temporal context enhanced referring video object segmentation
Xiao Hu, Basavaraj Hampiholi, Heiko Neumann, and Jochen Lang · 2024
Closest in time.
Phrase decoupling cross-modal hierarchical matching and progressive position correction for visual grounding
Minghong Xie, Mengzhao Wang, Huafeng Li, Yafei Zhang, Dapeng Tao, and Zhengtao Yu · 2025
Closest in time.
Graph-based referring expression comprehension with expression-guided selective filtering and noun-oriented reasoning
Jingcheng Ke, Qi Zhang, Jia Wang, Hongqing Ding, Pengfei Zhang, and Jie Wen · 2025
Closest in time.
Language prompt for autonomous driving
Dongming Wu, Wencheng Han, Yingfei Liu, Tiancai Wang, Cheng-zhong Xu, Xiangyu Zhang, and Jianbing Shen · 2025
Closest in time.
Cognitive disentanglement for referring multi-object tracking
Shaofeng Liang, Runwei Guan, Wangwang Lian, Daizong Liu, Xiaolou Sun, Dongming Wu, Yutao Yue, Weiping Ding, and Hui Xiong · 2025
Closest in time.
Cross-view referring multi-object tracking
Sijia Chen, En Yu, and Wenbing Tao · 2025
Closest in time.
Hff-tracker: A hierarchical fine-grained fusion tracker for referring multi-object tracking
Zeyong Zhao, Yanchao Hao, Minghao Zhang, Qingbin Liu, Bo Li, Dianbo Sui, Shizhu He, and Xi Chen · 2025
Closest in time.
Few-shot referring video single-and multi-object segmentation via cross-modal affinity with instance sequence matching
Heng Liu, Guanghui Li, Mingqi Gao, Xiantong Zhen, Feng Zheng, and Yang Wang · 2025
Closest in time.
Refergpt: Towards zero-shot referring multi-object tracking
Tzoulio Chamiti, Leandro Di Bella, Adrian Munteanu, and Nikos Deligiannis · 2025
Closest in time.
Nugrounding: A multi-view 3d visual grounding framework in autonomous driving
Fuhao Li, Huan Jin, Bin Gao, Liaoyuan Fan, Lihui Jiang, and Long Zeng · 2025
Closest in time.
Visual-linguistic feature alignment with semantic and kinematic guidance for referring multi-object tracking
Yizhe Li, Sanping Zhou, Zheng Qin, and Le Wang · 2025
Closest in time.
Refav: Towards planning-centric scenario mining
Cainan Davidson, Deva Ramanan, and Neehar Peri · 2025
Closest in time.
Cpany: Couple with any encoder to refer multi-object tracking
Weize Li, Yunhao Du, Qixiang Yin, Zhicheng Zhao, Fei Su, and Daqi Liu · 2025
Closest in time.