Fetching the paper…
Reading the bibliography…
A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt.
nuScenes: A multimodal dataset for autonomous driving
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2019 · 1903
Earlier work this paper cites.
Weakly-supervised spatio-temporally grounding natural sentence in video
Chen, Z.; Ma, L.; Luo, W.; and Wong, K.-Y. K. 2019 · 1906
Earlier work this paper cites.
Talk2car: Taking control of your self-driving car
Deruyttere, T.; Vandenhende, S.; Grujicic, D.; Van Gool, L.; and Moens, M.-F. 2019 · 1909
Earlier work this paper cites.
Evaluating multiple object tracking performance: the clear mot metrics
Bernardin, K.; and Stiefelhagen, R. 2008 · 2008
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Geiger, A.; Lenz, P.; and Urtasun, R. 2012 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Modeling context in referring expressions
Yu, L.; Poirson, P.; Yang, S.; Berg, A. C.; and Berg, T. L. 2016 · 2016
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Kim, J.; Rohrbach, A.; Darrell, T.; Canny, J.; and Akata, Z. 2018 · 2018
Earlier work this paper cites.
Object referring in videos with language and human gaze
Vasudevan, A. B.; Dai, D.; and Van Gool, L. 2018 · 2018
Earlier work this paper cites.
Video object segmentation with language referring expressions
Khoreva, A.; Rohrbach, A.; and Schiele, B. 2019 · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Earlier work this paper cites.
Tao: A large-scale benchmark for tracking any object
Dave, A.; Khurana, T.; Tokmakov, P.; Schmid, C.; and Ramanan, D. 2020 · 2020
Earlier work this paper cites.
Urvos: Unified referring video object segmentation network with a large-scale benchmark
Seo, S.; Lee, J.-Y.; and Han, B. 2020 · 2020
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S.; Sharma, P.; Ding, N.; and Soricut, R. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Katta, A.; Coombes, T.; Jitsev, J.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Center-based 3D Object Detection and Tracking
Yin, T.; Zhou, X.; and Krähenbühl, P. 2021 · 2021
Cited alongside, same era.
GRES: Generalized referring expression segmentation
Liu, C.; Ding, H.; and Jiang, X. 2023 · 2023
Closest in time.
DRAMA: Joint Risk Localization and Captioning in Driving
Malla, S.; Choi, C.; Dwivedi, I.; Choi, J. H.; and Li, J. 2023 · 2023
Closest in time.
Type-to-Track: Retrieve Any Object via Prompt-based Tracking
Nguyen, P.; Quach, K. G.; Kitani, K.; and Luu, K. 2023 · 2023
Closest in time.
https://chat.openai.com
OpenAI. 2023 · 2023
Closest in time.
Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking
Pang, Z.; Li, J.; Tokmakov, P.; Chen, D.; Zagoruyko, S.; and Wang, Y.-X. 2023 · 2023
Closest in time.
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Petr: Position embedding transformation for multi-view 3d object detection
Liu, Y.; Wang, T.; Zhang, X.; and Sun, J. 2022 · 2022
Cited alongside, same era.
Detr3d: 3d object detection from multi-view images via 3d-to-2d queries
Wang, Y.; Guizilini, V. C.; Zhang, T.; Wang, Y.; Zhao, H.; and Solomon, J. 2022 · 2022
Cited alongside, same era.
Motr: End-to-end multiple-object tracking with transformer
Zeng, F.; Dong, B.; Zhang, Y.; Wang, T.; Zhang, X.; and Wei, Y. 2022 · 2022
Cited alongside, same era.
End-to-end Autonomous Driving: Challenges and Frontiers
Chen, L.; Wu, P.; Chitta, K.; Jaeger, B.; Geiger, A.; and Li, H. 2023 · 2023
Cited alongside, same era.
ViP3D: End-to-end visual trajectory prediction via 3d agent queries
Gu, J.; Hu, C.; Zhang, T.; Chen, X.; Wang, Y.; Wang, Y.; and Zhao, H. 2023 · 2023
Cited alongside, same era.
Visual programming: Compositional visual reasoning without training
Gupta, T.; and Kembhavi, A. 2023 · 2023
Cited alongside, same era.
Planning-oriented autonomous driving
Hu, Y.; Yang, J.; Chen, L.; Li, K.; Sima, C.; Zhu, X.; Chai, S.; Du, S.; Lin, T.; Wang, W.; et al. 2023 · 2023
Cited alongside, same era.
Qian, T.; Chen, J.; Zhuo, L.; Jiao, Y.; and Jiang, Y.-G. 2023 · 2023
Closest in time.
Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection
Wang, S.; Liu, Y.; Wang, T.; Li, Y.; and Zhang, X. 2023 · 2023
Closest in time.
The 1st-place solution for cvpr 2023 openlane topology in autonomous driving challenge
Wu, D.; Jia, F.; Chang, J.; Li, Z.; Sun, J.; Han, C.; Li, S.; Liu, Y.; Ge, Z.; and Wang, T. 2023c · 2023
Closest in time.
Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
Bai, Y.; Wu, D.; Liu, Y.; Jia, F.; Mao, W.; Zhang, Z.; Zhao, Y.; Shen, J.; Wei, X.; Wang, T.; et al. 2024 · 2024
Closest in time.
Talk2BEV: Language-enhanced Bird’s-eye View Maps for Autonomous Driving
Dewangan, V.; Choudhary, T.; Chandhok, S.; Priyadarshan, S.; Jain, A.; Singh, A. K.; Srivastava, S.; Jatavallabhula, K. M.; and Krishna, K. M. 2024 · 2024
Closest in time.
ADA-Track: End-to-End Multi-Camera 3D Multi-Object Tracking with Alternating Detection and Association
Ding, S.; Schneider, L.; Cordts, M.; and Gall, J. 2024 · 2024
Closest in time.
Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning
Sachdeva, E.; Agarwal, N.; Chundi, S.; Roelofs, S.; Li, J.; Kochenderfer, M.; Choi, C.; and Dariush, B. 2024 · 2024
Closest in time.
Drivelm: Driving with graph visual question answering
Sima, C.; Renz, K.; Chitta, K.; Chen, L.; Zhang, H.; Xie, C.; Luo, P.; Geiger, A.; and Li, H. 2024 · 2024
Closest in time.