Fetching the paper…
Reading the bibliography…
City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures.
The ISPRS benchmark on urban object classification and 3D building reconstruction
F. Rottensteiner, G. Sohn, J. Jung, M. Gerke, C. Baillard, S. Benitez, and U. Breitkopf · 2012
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Paris-rue-madame database: a 3d mobile laser scanner dataset for benchmarking urban detection, segmentation and classification methods
A. Serna, B. Marcotegui, F. Goulette, and J.-E. Deschaud · 2014
Earlier work this paper cites.
Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
S. Song, S. P. Lichtenberg, and J. Xiao · 2015
Earlier work this paper cites.
Terramobilita/iqmulus urban point cloud analysis benchmark
B. Vallet, M. Brédif, A. Serna, B. Marcotegui, and N. Paparoditis · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Generation and Comprehension of Unambiguous Object Descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Potree: Rendering large point clouds in web browsers
M. Schütz et al · 2016
Earlier work this paper cites.
Semantic instance annotation of street scenes by 3d to 2d label transfer
J. Xie, M. Kiefel, M.-T. Sun, and A. Geiger · 2016
Earlier work this paper cites.
Modeling Context in Referring Expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Joint 2d-3d-semantic data for indoor scene understanding
I. Armeni, S. Sax, A. R. Zamir, and S. Savarese · 2017
Earlier work this paper cites.
Matterport3D: Learning from RGB-D data in indoor environments
A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang · 2017
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Semantic3D.Net: A new large-scale point cloud classification benchmark
T. Hackel, N. Savinov, L. Ladicky, J. D. Wegner, K. Schindler, and M. Pollefeys · 2017
Earlier work this paper cites.
3d semantic segmentation with submanifold sparse convolutional networks
B. Graham, M. Engelcke, and L. van der Maaten · 2018
Earlier work this paper cites.
Paris-Lille-3D: A large and high-quality ground-truth urban point cloud dataset for automatic segmentation and classification
X. Roynard, J.-E. Deschaud, and F. Goulette · 2018
Earlier work this paper cites.
Gibson Env: real-world perception for embodied agents
F. Xia, A. R. Zamir, Z.-Y. He, A. Sax, J. Malik, and S. Savarese · 2018
Earlier work this paper cites.
3d semantic segmentation with submanifold sparse convolutional networks
B. Graham, M. Engelcke, and L. van der Maaten · 2018
Earlier work this paper cites.
SemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences
J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, C. Stachniss, and J. Gall · 2019
Earlier work this paper cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
H. Chen, A. Suhr, D. Misra, N. Snavely, and Y. Artzi · 2019
Earlier work this paper cites.
Sun-spot: An rgb-d dataset with spatial referring expressions
C. Mauceri, M. Palmer, and C. Heckman · 2019
Earlier work this paper cites.
Rio: 3d object instance re-localization in changing indoor environments
J. Wald, A. Avetisyan, N. Navab, F. Tombari, and M. Niessner · 2019
Cited alongside, same era.
DublinCity: Annotated LiDAR point cloud and its applications
S. Zolanvari, S. Ruano, A. Rana, A. Cummins, R. E. da Silva, M. Rahbar, and A. Smolic · 2019
Cited alongside, same era.
ReferIt3D: Neural listeners for fine-grained 3d object identification in real-world scenes
P. Achlioptas, A. Abdelreheem, F. Xia, M. Elhoseiny, and L. J. Guibas · 2020
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Cited alongside, same era.
Campus3D: A photogrammetry point cloud benchmark for hierarchical understanding of outdoor scene
X. Li, C. Li, Z. Tong, A. Lim, J. Yuan, Y. Wu, J. Tang, and R. Huang · 2020
Cited alongside, same era.
D3net: A speaker-listener architecture for semi-supervised dense captioning and visual grounding in rgb-d scans
D. Z. Chen, Q. Wu, M. Nießner, and A. X. Chang · 2022
Later among the works it cites.
3dvqa: Visual question answering for 3d environments
Y. Etesam, L. Kochiev, and A. X. Chang · 2022
Later among the works it cites.
Sensaturban: Learning semantics from urban-scale photogrammetric point clouds
Q. Hu, B. Yang, S. Khalid, W. Xiao, N. Trigoni, and A. Markham · 2022
Later among the works it cites.
Multi-view transformer for 3d visual grounding
S. Huang, Y. Chen, J. Jia, and L. Wang · 2022
Later among the works it cites.
Bottom up top down detection transformers for language grounding in images and point clouds
A. Jain, N. Gkanatsios, I. Mediratta, and K. Fragkiadaki · 2022
Later among the works it cites.
Text2pos: Text-to-point-cloud cross-modal localization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Qi, Q. Wu, P. Anderson, X. Wang, W. Y. Wang, C. Shen, and A. van den Hengel · 2020
Cited alongside, same era.
Toronto-3D: A large-scale mobile LiDAR dataset for semantic segmentation of urban roadways
W. Tan, N. Qin, L. Ma, Y. Li, J. Du, G. Cai, K. Yang, and J. Li · 2020
Cited alongside, same era.
DALES: A large-scale aerial LiDAR data set for semantic segmentation
N. Varney, V. K. Asari, and Q. Graehling · 2020
Cited alongside, same era.
LASDU: A large-scale aerial LiDAR dataset for semantic labeling in dense urban areas
Z. Ye, Y. Xu, R. Huang, X. Tong, X. Li, X. Liu, K. Luan, L. Hoegner, and U. Stilla · 2020
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Cited alongside, same era.
ARKitscenes - a diverse real-world dataset for 3d indoor scene understanding using mobile RGB-d data
G. Baruch, Z. Chen, A. Dehghan, T. Dimry, Y. Feigin, P. Fu, T. Gebauer, B. Joffe, D. Kurz, A. Schwartz, and E. Shulman · 2021
Cited alongside, same era.
Scan2cap: Context-aware dense captioning in rgb-d scans
Z. Chen, A. Gholami, M. Nießner, and A. X. Chang · 2021
Cited alongside, same era.
M. Kolmet, Q. Zhou, A. Ošep, and L. Leal-Taixé · 2022
Later among the works it cites.
Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d
Y. Liao, J. Xie, and A. Geiger · 2022
Later among the works it cites.
Languagerefer: Spatial-language model for 3d visual grounding
J. Roh, K. Desingh, A. Farhadi, and D. Fox · 2022
Later among the works it cites.
Visual grounding in remote sensing images
Y. Sun, S. Feng, X. Li, Y. Ye, J. Kang, and X. Huang · 2022
Later among the works it cites.
Softgroup for 3d instance segmentation on 3d point clouds
T. Vu, K. Kim, T. M. Luu, X. T. Nguyen, and C. D. Yoo · 2022
Later among the works it cites.
Spatiality-guided transformer for 3d dense captioning on point clouds
H. Wang, C. Zhang, J. Yu, and W. Cai · 2022
Later among the works it cites.
Toward explainable and fine-grained 3d grounding through referring textual phrases
Z. Yuan, X. Yan, Z. Li, X. Li, Y. Guo, S. Cui, and Z. Li · 2022
Later among the works it cites.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense captioning
Z. Yuan, X. Yan, Y. Liao, Y. Guo, G. Li, S. Cui, and Z. Li · 2022
Later among the works it cites.
Towards explainable 3d grounded visual question answering: A new benchmark and strong baseline
L. Zhao, D. Cai, J. Zhang, L. Sheng, D. Xu, R. Zheng, Y. Zhao, L. Wang, and X. Fan · 2022
Later among the works it cites.
Sensaturban: Learning semantics from urban-scale photogrammetric point clouds
Q. Hu, B. Yang, S. Khalid, W. Xiao, N. Trigoni, and A. Markham · 2022
Later among the works it cites.
Softgroup for 3d instance segmentation on 3d point clouds
T. Vu, K. Kim, T. M. Luu, X. T. Nguyen, and C. D. Yoo · 2022
Later among the works it cites.
STPLS3D: A Large-Scale Synthetic and Real Aerial Photogrammetry 3D Point Cloud Dataset
M. Chen, Q. Hu, Z. Yu, H. Thomas, A. Feng, Yu. Hou, K. McCullough, F. Ren, L. Soibelman · 2022
Later among the works it cites.
Sqa3d: Situated question answering in 3d scenes
X. Ma, S. Yong, Z. Zheng, Q. Li, Y. Liang, S.-C. Zhu, and S. Huang · 2023
Closest in time.
3d change localization and captioning from dynamic scans of indoor scenes
Y. Qiu, S. Yamamoto, R. Yamada, R. Suzuki, H. Kataoka, K. Iwata, and Y. Satoh · 2023
Closest in time.
Eda: Explicit text-decoupling and dense alignment for 3d visual grounding
Y. Wu, X. Cheng, R. Zhang, Z. Cheng, and J. Zhang · 2023
Closest in time.
Rsvg: Exploring data and models for visual grounding on remote sensing data
Y. Zhan, Z. Xiong, and Y. Yuan · 2023
Closest in time.
Scanents3d: Exploiting phrase-to-3d-object correspondences for improved visio-linguistic models in 3d scenes
A. Abdelreheem, K. Olszewski, H.-Y. Lee, P. Wonka, and P. Achlioptas · 2024
Closest in time.