Fetching the paper…
Reading the bibliography…
Masked Autoencoders learn strong visual representations and achieve state-of-the-art results in several independent modalities, yet very few works have addressed their capabilities in multi-modality settings.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Deep sliding shapes for amodal 3d object detection in rgb-d images, 2015
Shuran Song and Jianxiong Xiao · 2015
Earlier work this paper cites.
A point set generation network for 3d object reconstruction from a single image, 2016
Haoqiang Fan, Hao Su, and Leonidas Guibas · 2016
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, koray kavukcuoglu, and Daan Wierstra · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks, 2017
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Joint 3d proposal generation and object detection from view aggregation, 2017
Jason Ku, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven Waslander · 2017
Earlier work this paper cites.
2d-driven 3d object detection in rgb-d images
Jean Lahoud and Bernard Ghanem · 2017
Earlier work this paper cites.
Frustum pointnets for 3d object detection from rgb-d data, 2017
Charles R. Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J. Guibas · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning, 2017
Jake Snell, Kevin Swersky, and Richard S. Zemel · 2017
Earlier work this paper cites.
Learning to compare: Relation network for few-shot learning, 2017
Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip H. S. Torr, and Timothy M. Hospedales · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Pointfusion: Deep sensor fusion for 3d bounding box estimation, 2017
Danfei Xu, Dragomir Anguelov, and Ashesh Jain · 2017
Earlier work this paper cites.
Voxelnet: End-to-end learning for point cloud based 3d object detection, 2017
Yin Zhou and Oncel Tuzel · 2017
Earlier work this paper cites.
3d-sis: 3d semantic instance segmentation of rgb-d scans, 2018
Ji Hou, Angela Dai, and Matthias Nießner · 2018
Earlier work this paper cites.
Tadam: Task dependent adaptive metric for improved few-shot learning
Boris N. Oreshkin, Pau Rodriguez, and Alexandre Lacoste · 2018
Earlier work this paper cites.
Multi-level fusion based 3d object detection from monocular images
Bin Xu and Zhenzhong Chen · 2018
Earlier work this paper cites.
Meta-learning with differentiable closed-form solvers
Luca Bertinetto, Joao F. Henriques, Philip Torr, and Andrea Vedaldi · 2019
Earlier work this paper cites.
Deep hough voting for 3d object detection in point clouds, 2019
Charles R. Qi, Or Litany, Kaiming He, and Leonidas J. Guibas · 2019
Earlier work this paper cites.
3d object proposal generation and detection from point cloud
S Shi, X Wang, H Pointrcnn Li, et al · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Debiased contrastive learning
Ching-Yao Chuang, Joshua Robinson, Yen-Chen Lin, Antonio Torralba, and Stefanie Jegelka · 2020
Cited alongside, same era.
Generative sparse detection networks for 3d single-shot object detection
JunYoung Gwak, Christopher B Choy, and Silvio Savarese · 2020
Cited alongside, same era.
Exploring data-efficient 3d scene understanding with contrastive scene contexts, 2020
Ji Hou, Benjamin Graham, Matthias Nießner, and Saining Xie · 2020
Cited alongside, same era.
Deep continuous fusion for multi-sensor 3d object detection, 2020
Ming Liang, Bin Yang, Shenlong Wang, and Raquel Urtasun · 2020
Cited alongside, same era.
P4contrast: Contrastive learning with pairs of point-pixel pairs for RGB-D scene understanding, 2020
Yunze Liu, Li Yi, Shanghang Zhang, Qingnan Fan, Thomas A. Funkhouser, and Hao Dong · 2020
Cited alongside, same era.
Tanet: Robust 3d object detection from point clouds with triple attention
Graph debiased contrastive learning with joint representation clustering
Han Zhao, Xu Yang, Zhenru Wang, Erkun Yang, and Cheng Deng · 2021
Later among the works it cites.
Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding
Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri, Kanchana Thilakarathna, and Ranga Rodrigo · 2022
Later among the works it cites.
MultiMAE: Multi-modal multi-task masked autoencoders
Roman Bachmann, David Mizrahi, Andrei Atanov, and Amir Zamir · 2022
Later among the works it cites.
Pos-bert: Point cloud one-stage bert pre-training
Kexue Fu, Peng Gao, ShaoLei Liu, Renrui Zhang, Yu Qiao, and Manning Wang · 2022
Later among the works it cites.
Convmae: Masked convolution meets masked autoencoders
Peng Gao, Teli Ma, Hongsheng Li, Jifeng Dai, and Yu Qiao · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhe Liu, Xin Zhao, Tengteng Huang, Ruolan Hu, Yu Zhou, and Xiang Bai · 2020
Cited alongside, same era.
Pv-rcnn: Point-voxel feature set abstraction for 3d object detection
Shaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li · 2020
Cited alongside, same era.
Pointcontrast: Unsupervised pre-training for 3d point cloud understanding
Saining Xie, Jiatao Gu, Demi Guo, Charles R Qi, Leonidas Guibas, and Or Litany · 2020
Cited alongside, same era.
H3dnet: 3d object detection using hybrid geometric primitives, 2020
Zaiwei Zhang, Bo Sun, Haitao Yang, and Qixing Huang · 2020
Cited alongside, same era.
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H. S. Torr, and Vladlen Koltun · 2020
Cited alongside, same era.
End-to-end object detection with adaptive clustering transformer
Minghang Zheng, Peng Gao, Renrui Zhang, Kunchang Li, Xiaogang Wang, Hongsheng Li, and Hao Dong · 2020
Cited alongside, same era.
BEiT: BERT pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei · 2021
Cited alongside, same era.
Later among the works it cites.
Multimodal masked autoencoders learn transferable representations
Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurmans, Sergey Levine, and Pieter Abbeel · 2022
Later among the works it cites.
Calip: Zero-shot enhancement of clip with parameter-free attention
Ziyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma, Xupeng Miao, Xuming He, and Bin Cui · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Masked autoencoders for self-supervised learning on automotive point clouds
Georg Hess, Johan Jaxing, Elias Svensson, David Hagerman, Christoffer Petersson, and Lennart Svensson · 2022
Later among the works it cites.
Tig-bev: Multi-view bev 3d object detection via target inner-geometry learning
Peixiang Huang, Li Liu, Renrui Zhang, Song Zhang, Xinli Xu, Baichao Wang, and Guoyi Liu · 2022
Later among the works it cites.
Cross-modal learning for domain adaptation in 3d semantic segmentation
Maximilian Jaritz, Tuan-Hung Vu, Raoul De Charette, Émilie Wirbel, and Patrick Pérez · 2022
Later among the works it cites.
Simipu: Simple 2d image and 3d point cloud unsupervised pre-training for spatial-aware visual representations
Zhenyu Li, Zehui Chen, Ang Li, Liangji Fang, Qinhong Jiang, Xianming Liu, Junjun Jiang, Bolei Zhou, and Hang Zhao · 2022
Later among the works it cites.
Voxel-mae: Masked autoencoders for pre-training large-scale point clouds
Chen Min, Dawei Zhao, Liang Xiao, Yiming Nie, and Bin Dai · 2022
Later among the works it cites.
Masked autoencoders for point cloud self-supervised learning
Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan · 2022
Later among the works it cites.
Eda: Explicit text-decoupling and dense alignment for 3d visual and language learning
Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng, and Jian Zhang · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2022
Later among the works it cites.
Boosting 3d object detection via object-focused image fusion
Hao Yang, Chen Shi, Yihong Chen, and Liwei Wang · 2022
Later among the works it cites.
i-mae: Are latent representations in masked autoencoders linearly separable?
Kevin Zhang and Zhiqiang Shen · 2022
Later among the works it cites.
Point-m2ae: Multi-scale masked autoencoders for hierarchical point cloud pre-training
Renrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang, Bin Zhao, Dong Wang, Yu Qiao, and Hongsheng Li · 2022
Later among the works it cites.
Pointclip: Point cloud understanding by clip
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li · 2022
Later among the works it cites.
Monodetr: Depth-aware transformer for monocular 3d object detection
Renrui Zhang, Han Qiu, Tai Wang, Xuanzhuo Xu, Ziyu Guo, Yu Qiao, Peng Gao, and Hongsheng Li · 2022
Later among the works it cites.
Learning 3d representations from 2d pre-trained models via image-to-point masked autoencoders
Renrui Zhang, Liuhui Wang, Yu Qiao, Peng Gao, and Hongsheng Li · 2022
Later among the works it cites.
Can language understand depth?
Renrui Zhang, Ziyao Zeng, Ziyu Guo, and Yafeng Li · 2022
Later among the works it cites.
Pointclip v2: Adapting clip for powerful 3d open-world learning
Xiangyang Zhu, Renrui Zhang, Bowei He, Ziyao Zeng, Shanghang Zhang, and Peng Gao · 2022
Later among the works it cites.
Parameter is not all you need: Starting from non-parametric networks for 3d point cloud analysis
Renrui Zhang, Liuhui Wang, Yali Wang, Peng Gao, Hongsheng Li, and Jianbo Shi · 2023
Closest in time.