Fetching the paper…
Reading the bibliography…
Performing 3D dense captioning and visual grounding requires a common and shared understanding of the underlying multimodal relationships.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
ShapeNet: An information-rich 3D model repository
Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
ScanNet: Richly-annotated 3D reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
3DMV: Joint 3D-multi-view prediction for 3D semantic scene segmentation
Angela Dai and Matthias Nießner · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Frustum pointnets for 3D object detection from RGB-D data
Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
4D spatio-temporal convnets: Minkowski convolutional neural networks
Christopher Choy, JunYoung Gwak, and Silvio Savarese · 2019
Earlier work this paper cites.
3D-SIS: 3D semantic instance segmentation of RGB-D scans
Ji Hou, Angela Dai, and Matthias Nießner · 2019
Earlier work this paper cites.
3D instance segmentation via multi-task metric learning
Jean Lahoud, Bernard Ghanem, Marc Pollefeys, and Martin R Oswald · 2019
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Deep Hough voting for 3D object detection in point clouds
Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Cited alongside, same era.
ReferIt3D: Neural listeners for fine-grained 3D object identification in real-world scenes
Panos Achlioptas, Ahmed Abdelreheem, Fei Xia, Mohamed Elhoseiny, and Leonidas Guibas · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
SimVLM: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao · 2021
Later among the works it cites.
Image2point: 3D point-cloud understanding with pretrained 2D convnets
Chenfeng Xu, Shijia Yang, Bohan Zhai, Bichen Wu, Xiangyu Yue, Wei Zhan, Peter Vajda, Kurt Keutzer, and Masayoshi Tomizuka · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ScanRefer: 3D object localization in RGB-D scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner · 2020
Cited alongside, same era.
3D-MPA: Multi-proposal aggregation for 3D semantic instance segmentation
Francis Engelmann, Martin Bokeloh, Alireza Fathi, Bastian Leibe, and Matthias Nießner · 2020
Cited alongside, same era.
RevealNet: Seeing behind objects in RGB-D scans
Ji Hou, Angela Dai, and Matthias Nießner · 2020
Cited alongside, same era.
PointGroup: Dual-set point grouping for 3D instance segmentation
Li Jiang, Hengshuang Zhao, Shaoshuai Shi, Shu Liu, Chi-Wing Fu, and Jiaya Jia · 2020
Cited alongside, same era.
ImVoteNet: Boosting 3D object detection in point clouds with image votes
Charles R Qi, Xinlei Chen, Or Litany, and Leonidas J Guibas · 2020
Cited alongside, same era.
Unified vision-language pre-training for image captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao · 2020
Cited alongside, same era.
Back-tracing representative points for voting-based 3D object detection in point clouds
Bowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang, and Dong Xu · 2021
Cited alongside, same era.
Zhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang, Sheng Wang, Zhen Li, and Shuguang Cui · 2021
Later among the works it cites.
3DVG-Transformer: Relation modeling for visual grounding on point clouds
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu · 2021
Later among the works it cites.
3DJCG: A unified framework for joint dense captioning and visual grounding on 3D point clouds
Daigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng, and Dong Xu · 2022
Closest in time.
Semantic abstraction: Open-world 3D scene understanding from 2D vision-language models
Huy Ha and Shuran Song · 2022
Closest in time.
Clip2point: Transfer clip to point cloud classification with image-depth pre-training
Tianyu Huang, Bowen Dong, Yunhan Yang, Xiaoshui Huang, Rynson WH Lau, Wanli Ouyang, and Wangmeng Zuo · 2022
Closest in time.
Bottom up top down detection transformers for language grounding in images and point clouds
Ayush Jain, Nikolaos Gkanatsios, Ishita Mediratta, and Katerina Fragkiadaki · 2022
Closest in time.
More: Multi-order relation mining for dense captioning in 3d scenes
Yang Jiao, Shaoxiang Chen, Zequn Jie, Jingjing Chen, Lin Ma, and Yu-Gang Jiang · 2022
Closest in time.
3d-sps: Single-stage 3d visual grounding via referred point progressive selection
Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao, Haibing Ren, Hao Shen, Huaxia Xia, and Si Liu · 2022
Closest in time.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Closest in time.
X-trans2cap: Cross-modal knowledge transfer using transformer for 3D dense captioning
Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Shuguang Cui, and Zhen Li · 2022
Closest in time.
PointCLIP: Point cloud understanding by CLIP
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li · 2022
Closest in time.