Fetching the paper…
Reading the bibliography…
We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019c · 1907
Earlier work this paper cites.
Improving referring expression grounding with cross-modal attention-guided erasing
Liu, X.; Wang, Z.; Shao, J.; Wang, X.; and Li, H. 2019b · 1959
Earlier work this paper cites.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Wang, P.; Wu, Q.; Cao, J.; Shen, C.; Gao, L.; and Hengel, A. v. d. 2019 · 1968
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Geiger, A.; Lenz, P.; and Urtasun, R. 2012 · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Query-guided regression network with context policy for phrase grounding
Chen, K.; Kovvuri, R.; and Nevatia, R. 2017 · 2017
Earlier work this paper cites.
Modeling relationships in referential expressions with compositional modular networks
Hu, R.; Rohrbach, M.; Andreas, J.; Darrell, T.; and Saenko, K. 2017 · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017 · 2017
Earlier work this paper cites.
Real-time referring expression comprehension by single-stage grounding network
Chen, X.; Ma, L.; Chen, J.; Jie, Z.; Liu, W.; and Luo, J. 2018 · 2018
Earlier work this paper cites.
Grounding referring expressions in images by variational context
Zhang, H.; Niu, Y.; and Chang, S.-F. 2018 · 2018
Earlier work this paper cites.
M3d-rpn: Monocular 3d region proposal network for object detection
Brazil, G.; and Liu, X. 2019 · 2019
Earlier work this paper cites.
Sun-spot: An rgb-d dataset with spatial referring expressions
Mauceri, C.; Palmer, M.; and Heckman, C. 2019 · 2019
Earlier work this paper cites.
Monogrnet: A geometric reasoning network for monocular 3d object localization
Qin, Z.; Wang, J.; and Lu, Y. 2019 · 2019
Earlier work this paper cites.
Generalized intersection over union: A metric and a loss for bounding box regression
Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; and Savarese, S. 2019 · 2019
Earlier work this paper cites.
Zero-shot grounding of objects from natural language queries
Sadhu, A.; Chen, K.; and Nevatia, R. 2019 · 2019
Earlier work this paper cites.
Dynamic graph attention for referring expression comprehension
Yang, S.; Li, G.; and Yu, Y. 2019 · 2019
Earlier work this paper cites.
A fast and accurate one-stage approach to visual grounding
Yang, Z.; Gong, B.; Wang, L.; Huang, W.; Yu, D.; and Luo, J. 2019 · 2019
Earlier work this paper cites.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Achlioptas, P.; Abdelreheem, A.; Xia, F.; Elhoseiny, M.; and Guibas, L. 2020 · 2020
Earlier work this paper cites.
MonoFENet: Monocular 3D Object Detection With Feature Enhancement Networks
Bao, W.; Xu, B.; and Chen, Z. 2020 · 2020
Cited alongside, same era.
Kinematic 3d object detection in monocular video
Brazil, G.; Pons-Moll, G.; Liu, X.; and Schiele, B. 2020 · 2020
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
Chen, D. Z.; Chang, A. X.; and Nießner, M. 2020 · 2020
Cited alongside, same era.
MonoPair: Monocular 3D Object Detection Using Pairwise Spatial Relationships
Chen, Y.; Tai, L.; Sun, K.; and Li, M. 2020 · 2020
Cited alongside, same era.
Learning depth-guided convolutions for monocular 3d object detection
Ding, M.; Huo, Y.; Yi, H.; Wang, Z.; Shi, J.; Lu, Z.; and Luo, P. 2020 · 2020
Cited alongside, same era.
Smoke: Single-stage monocular 3d object detection via keypoint estimation
Liu, Z.; Wu, Z.; and Tóth, R. 2020 · 2020
Sat: 2d semantics assisted training for 3d visual grounding
Yang, Z.; Zhang, S.; Wang, L.; and Luo, J. 2021 · 2021
Later among the works it cites.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Yuan, Z.; Yan, X.; Liao, Y.; Zhang, R.; Wang, S.; Li, Z.; and Cui, S. 2021 · 2021
Later among the works it cites.
Objects are different: Flexible monocular 3d object detection
Zhang, Y.; Lu, J.; and Zhou, J. 2021 · 2021
Later among the works it cites.
3DVG-Transformer: Relation modeling for visual grounding on point clouds
Zhao, L.; Cai, D.; Sheng, L.; and Xu, D. 2021 · 2021
Later among the works it cites.
3djcg: A unified framework for joint dense captioning and visual grounding on 3d point clouds
Cai, D.; Zhao, L.; Zhang, J.; Sheng, L.; and Xu, D. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reverie: Remote embodied visual referring expression in real indoor environments
Qi, Y.; Wu, Q.; Anderson, P.; Wang, X.; Wang, W. Y.; Shen, C.; and Hengel, A. v. d. 2020 · 2020
Cited alongside, same era.
Relationship-embedded representation learning for grounding referring expressions
Yang, S.; Li, G.; and Yu, Y. 2020 · 2020
Cited alongside, same era.
Improving one-stage visual grounding by recursive sub-query construction
Yang, Z.; Chen, T.; Wang, L.; and Luo, J. 2020 · 2020
Cited alongside, same era.
TransVG: End-to-End Visual Grounding With Transformers
Deng, J.; Yang, Z.; Chen, T.; Zhou, W.; and Li, H. 2021 · 2021
Cited alongside, same era.
Free-form description guided 3d visual graph network for object grounding in point cloud
Feng, M.; Li, Z.; Li, Q.; Zhang, L.; Zhang, X.; Zhu, G.; Zhang, H.; Wang, Y.; and Mian, A. 2021 · 2021
Cited alongside, same era.
Transrefer3d: Entity-and-relation aware transformer for fine-grained 3d visual grounding
He, D.; Zhao, Y.; Luo, J.; Hui, T.; Huang, S.; Zhang, A.; and Liu, S. 2021 · 2021
Cited alongside, same era.
D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding
Chen, D. Z.; Wu, Q.; Nießner, M.; and Chang, A. X. 2022 · 2022
Later among the works it cites.
Pseudo-stereo for monocular 3d object detection in autonomous driving
Chen, Y.-N.; Dai, H.; and Ding, Y. 2022 · 2022
Later among the works it cites.
Visual Grounding with Transformers
Du, Y.; Fu, Z.; Liu, Q.; and Wang, Y. 2022 · 2022
Later among the works it cites.
Learning to Compose and Reason with Language Tree Structures for Visual Grounding
Hong, R.; Liu, D.; Mo, X.; He, X.; and Zhang, H. 2022 · 2022
Later among the works it cites.
Progressive Language-Customized Visual Feature Learning for One-Stage Visual Grounding
Liao, Y.; Zhang, A.; Chen, Z.; Hui, T.; and Liu, S. 2022 · 2022
Later among the works it cites.
Languagerefer: Spatial-language model for 3d visual grounding
Roh, J.; Desingh, K.; Farhadi, A.; and Fox, D. 2022 · 2022
Later among the works it cites.
A proposal-free one-stage framework for referring expression comprehension and generation via dense cross-attention
Sun, M.; Suo, W.; Wang, P.; Zhang, Y.; and Wu, Q. 2022 · 2022
Later among the works it cites.
Improving Visual Grounding with Visual-Linguistic Verification and Iterative Reasoning
Yang, L.; Xu, Y.; Yuan, C.; Liu, W.; Li, B.; and Hu, W. 2022 · 2022
Later among the works it cites.
Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual Grounding
Ye, J.; Tian, J.; Yan, M.; Yang, X.; Wang, X.; Zhang, J.; He, L.; and Lin, X. 2022 · 2022
Later among the works it cites.
MonoDETR: depth-guided transformer for monocular 3D object detection
Zhang, R.; Qiu, H.; Wang, T.; Guo, Z.; Xu, X.; Qiao, Y.; Gao, P.; and Li, H. 2022 · 2022
Later among the works it cites.
Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild
Brazil, G.; Kumar, A.; Straub, J.; Ravi, N.; Johnson, J.; and Gkioxari, G. 2023 · 2023
Closest in time.
Lin, Z.; Peng, X.; Cong, P.; Hou, Y.; Zhu, X.; Yang, S.; and Ma, Y. 2023 · 2023
Closest in time.
RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data
Zhan, Y.; Xiong, Z.; and Yuan, Y. 2023 · 2023
Closest in time.