Fetching the paper…
Reading the bibliography…
Recent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the KITTI vision benchmark suite,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2012, pp. 3354–3361
2012
Earlier work this paper cites.
X. Chen, K. Kundu, Y. Zhu, A. G. Berneshawi, H. Ma, S. Fidler, and R. Urtasun, “3d object proposals for accurate object class detection,” in Proc. Adv. Neural Inf. Process. Syst. , 2015, pp. 424–432
2015
Earlier work this paper cites.
S. Song, S. P. Lichtenberg, and J. Xiao, “SUN RGB-D: a RGB-D scene understanding benchmark suite,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2015, pp. 567–576
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2016, pp. 770–778
2016
Earlier work this paper cites.
A. Mousavian, D. Anguelov, J. Flynn, and J. Kosecka, “3d bounding box estimation using deep learning and geometry,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2017, pp. 5632–5640
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
Z. Qin, J. Wang, and Y. Lu, “Monogrnet: A geometric reasoning network for monocular 3d object localization,” in Proc. AAAI Conf. Artif. Intell. , 2019, pp. 8851–8858
2019
Earlier work this paper cites.
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom, “Pointpillars: Fast encoders for object detection from point clouds,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2019, pp. 12 697–12 705
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Am. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol. , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the continuity of rotation representations in neural networks,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2019, pp. 5745–5753
2019
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in Proc. Eur. Conf. Comput. Vis. , vol. 12346, 2020, pp. 213–229
2020
Earlier work this paper cites.
P. Li, H. Zhao, P. Liu, and F. Cao, “Rtm3d: Real-time monocular 3d detection from object keypoints for autonomous driving,” in Proc. Eur. Conf. Comput. Vis. , vol. 12348, 2020, pp. 644–660
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2020, pp. 11 618–11 628
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proc. Int. Conf. Mach. Learn. , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
Y. Zhang, J. Lu, and J. Zhou, “Objects are different: Flexible monocular 3d object detection,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 3289–3298
2021
Earlier work this paper cites.
X. Shi, Q. Ye, X. Chen, C. Chen, Z. Chen, and T. Kim, “Geometry-based distance decomposition for monocular 3d object detection,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 15 152–15 161
2021
Earlier work this paper cites.
Y. Lu, X. Ma, L. Yang, T. Zhang, Y. Liu, Q. Chu, J. Yan, and W. Ouyang, “Geometry uncertainty projection network for monocular 3d object detection,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 3091–3101
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. Int. Conf. Learn. Represent. , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, Z. Yang, A. Chouard, P. Sun, J. Ngiam, V. Vasudevan, A. McCauley, J. Shlens, and D. Anguelov, “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. , 2021, pp. 9690–9699
2021
Cited alongside, same era.
A. Ahmadyan, L. Zhang, A. Ablavatski, J. Wei, and M. Grundmann, “Objectron: A large scale dataset of object-centric videos in the wild with pose annotations,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2021, pp. 7822–7831
2021
Cited alongside, same era.
T. Shen, D. Li, F.-Y. Wang, and H. Huang, “Depth-aware multi-person 3d pose estimation with multi-scale waterfall representations,” IEEE Transactions on Multimedia , 2022
2022
Cited alongside, same era.
G. Hua, H. Liu, W. Li, Q. Zhang, R. Ding, and X. Xu, “Weakly-supervised 3d human pose estimation with cross-view u-shaped graph convolutional network,” IEEE Transactions on Multimedia , 2022
2023
Later among the works it cites.
W. Wang, Z. Chen, X. Chen, J. Wu, X. Zhu, G. Zeng, P. Luo, T. Lu, J. Zhou, Y. Qiao, and J. Dai, “Visionllm: Large language model is also an open-ended decoder for vision-centric tasks,” in Proc. Adv. Neural Inf. Process. Syst. , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learn. Represent. , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 11 966–11 976
2022
Cited alongside, same era.
K.-C. Huang, T.-H. Wu, H.-T. Su, and W. H. Hsu, “Monodtr: Monocular 3d object detection with depth-aware transformer,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 4002–4011
2022
Cited alongside, same era.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in Proc. Int. Conf. Mach. Learn. , 2022, pp. 23 318–23 340
2022
Cited alongside, same era.
J. Chen, H. Guo, K. Yi, B. Li, and M. Elhoseiny, “Visualgpt: Data-efficient adaptation of pretrained language models for image captioning,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. , 2022, pp. 18 030–18 040
2022
Cited alongside, same era.
H. Zhang, J. Wang, J. Zhang, T. Zhang, and B. Zhong, “One-stream vision-language memory network for object tracking,” IEEE Transactions on Multimedia , 2023
2023
Cited alongside, same era.
P. An, Y. Duan, Y. Huang, J. Ma, Y. Chen, L. Wang, Y. Yang, and Q. Liu, “Sp-det: Leveraging saliency prediction for voxel-based 3d object detection in sparse point cloud,” IEEE Transactions on Multimedia , 2023
2023
Cited alongside, same era.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez et al. , “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” See https://vicuna. lmsys. org (accessed 14 April 2023) , vol. 2, no. 3, p. 6, 2023
2023
Later among the works it cites.
Y. Li, B. Hu, X. Chen, L. Ma, Y. Xu, and M. Zhang, “Lmeye: An interactive perception network for large language models,” IEEE Transactions on Multimedia , 2024
2024
Closest in time.
S. Zhao, H. Yao, C. Lin, Y. Gao, and G. Ding, “Multi-source-free domain adaptive object detection,” Int. J. Comput. Vis. , pp. 1–33, 2024
2024
Closest in time.
Y. Zhan, Y. Yuan, and Z. Xiong, “Mono3dvg: 3d visual grounding in monocular images,” in Proc. AAAI Conf. Artif. Intell. , 2024, pp. 6988–6996
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Zhang, L. Wu, Z. Zhang, T. Yu, C. Ma, X. Jin, X. Yang, and W. Zeng, “Unleash the power of vision-language models by visual attention prompt and multi-modal interaction,” IEEE Transactions on Multimedia , 2024
2024
Closest in time.
2024
Closest in time.
OpenAI, “Gpt-4o: The cutting-edge advancement in multimodal llm,” 2024. [Online]. Available: https://openai.com/index/hello-gpt-4o/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
F. Yang, H. Chen, Y. He, S. Zhao, C. Zhang, K. Ni, and G. Ding, “Geometry-guided domain generalization for monocular 3d object detection,” in Proc. AAAI Conf. Artif. Intell. , 2024, pp. 6467–6476
2024
Closest in time.
Q. Team, “Qwen2.5-vl,” January 2025. [Online]. Available: https://qwenlm.github.io/blog/qwen2.5-vl/
2025
Closest in time.