Fetching the paper…
Reading the bibliography…
The integration of data from diverse sensor modalities (e.g., camera and LiDAR) constitutes a prevalent methodology within the ambit of autonomous driving scenarios.
Sur une courbe, qui remplit toute une aire plane
G. Peano and G. Peano · 1990
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi, H. Su, K. Mo, and L. J. Guibas · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The best of both worlds: Combining recent advances in neural machine translation
M. X. Chen, O. Firat, A. Bapna, M. Johnson, W. Macherey, G. Foster, L. Jones, N. Parmar, M. Schuster, Z. Chen, et al · 2018
Earlier work this paper cites.
Second: Sparsely embedded convolutional detection
Y. Yan, Y. Mao, and B. Li · 2018
Earlier work this paper cites.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Y. Zhou and O. Tuzel · 2018
Earlier work this paper cites.
4d spatio-temporal convnets: Minkowski convolutional neural networks
C. Choy, J. Gwak, and S. Savarese · 2019
Earlier work this paper cites.
Pointpillars: Fast encoders for object detection from point clouds
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom · 2019
Earlier work this paper cites.
An energy and gpu-computation efficient backbone network for real-time object detection
Y. Lee, J.-w. Hwang, S. Lee, Y. Bae, and J. Park · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom · 2020
Earlier work this paper cites.
Using survival theory in early pattern detection for viral cascades
X. Gao, X. Jia, C. Yang, and G. Chen · 2020
Earlier work this paper cites.
Epnet: Enhancing point features with image semantics for 3d object detection
T. Huang, Z. Liu, X. Chen, and X. Bai · 2020
Earlier work this paper cites.
Sentimem: Attentive memory networks for sentiment classification in user review
X. Jia, Q. Wu, X. Gao, and G. Chen · 2020
Earlier work this paper cites.
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d
J. Philion and S. Fidler · 2020
Earlier work this paper cites.
Glu variants improve transformer, 2020
N. Shazeer · 2020
Earlier work this paper cites.
Openpcdet: An open-source toolbox for 3d object detection from point clouds
O. D. Team · 2020
Earlier work this paper cites.
Pointpainting: Sequential fusion for 3d object detection
S. Vora, A. H. Lang, B. Helou, and O. Beijbom · 2020
Earlier work this paper cites.
End-to-end multi-view fusion for 3d object detection in lidar point clouds
Y. Zhou, P. Sun, Y. Zhang, D. Anguelov, J. Gao, T. Ouyang, J. Guo, J. Ngiam, and V. Vasudevan · 2020
Earlier work this paper cites.
Deformable detr: Deformable transformers for end-to-end object detection
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai · 2020
Earlier work this paper cites.
Pct: Point cloud transformer
M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R. R. Martin, and S.-M. Hu · 2021
Earlier work this paper cites.
Bevdet: High-performance multi-camera 3d object detection in bird-eye-view
J. Huang, G. Huang, Z. Zhu, Y. Ye, and D. Du · 2021
Earlier work this paper cites.
Ide-net: Interactive driving event and pattern extraction from human data
X. Jia, L. Sun, M. Tomizuka, and W. Zhan · 2021
Earlier work this paper cites.
Ide-net: Interactive driving event and pattern extraction from human data
X. Jia, L. Sun, M. Tomizuka, and W. Zhan · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Earlier work this paper cites.
Pointaugmenting: Cross-modal augmentation for 3d object detection
C. Wang, C. Ma, M. Zhu, and X. Yang · 2021
Earlier work this paper cites.
Fcos3d: Fully convolutional one-stage monocular 3d object detection
T. Wang, X. Zhu, J. Pang, and D. Lin · 2021
Earlier work this paper cites.
Center-based 3d object detection and tracking
T. Yin, X. Zhou, and P. Krahenbuhl · 2021
Earlier work this paper cites.
Multimodal virtual point 3d detection
T. Yin, X. Zhou, and P. Krähenbühl · 2021
Earlier work this paper cites.
Transfusion: Robust lidar-camera fusion for 3d object detection with transformers
X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai · 2022
Cited alongside, same era.
Spconv: Spatially sparse convolution library
S. Contributors · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Embracing single stride 3d object detector with sparse transformer
L. Fan, Z. Pang, T. Zhang, Y.-X. Wang, H. Zhao, F. Wang, N. Wang, and Z. Zhang · 2022
Cited alongside, same era.
Bevdet4d: Exploit temporal cues in multi-camera 3d object detection
J. Huang and G. Huang · 2022
Cited alongside, same era.
Multi-agent trajectory prediction by combining egocentric and allocentric views
Dsvt: Dynamic sparse voxel transformer with rotated sets
H. Wang, C. Shi, S. Shi, M. Lei, S. Wang, D. He, B. Schiele, and L. Wang · 2023
Later among the works it cites.
Unitr: A unified and efficient multi-modal transformer for bird’s-eye-view representation
H. Wang, H. Tang, S. Shi, A. Li, Z. Li, B. Schiele, and L. Wang · 2023
Later among the works it cites.
Exploring object-centric temporal modeling for efficient multi-view 3d object detection
S. Wang, Y. Liu, T. Wang, Y. Li, and X. Zhang · 2023
Later among the works it cites.
Policy pre-training for autonomous driving via self-supervised geometric modeling
P. Wu, L. Chen, H. Li, X. Jia, J. Yan, and Y. Qiao · 2023
Later among the works it cites.
Policy pre-training for autonomous driving via self-supervised geometric modeling
P. Wu, L. Chen, H. Li, X. Jia, J. Yan, and Y. Qiao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Jia, L. Sun, H. Zhao, M. Tomizuka, and W. Zhan · 2022
Cited alongside, same era.
Unifying voxel-based representation with transformer for 3d object detection
Y. Li, Y. Chen, X. Qi, Z. Li, J. Sun, and J. Jia · 2022
Cited alongside, same era.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai · 2022
Cited alongside, same era.
Cbnet: A composite backbone network architecture for object detection
T. Liang, X. Chu, Y. Liu, Y. Wang, Z. Tang, W. Chu, J. Chen, and H. Ling · 2022
Cited alongside, same era.
Bevfusion: A simple and robust lidar-camera fusion framework
T. Liang, H. Xie, K. Yu, Z. Xia, Z. Lin, Y. Wang, T. Tang, B. Wang, and Z. Tang · 2022
Cited alongside, same era.
Petr: Position embedding transformation for multi-view 3d object detection
Y. Liu, T. Wang, X. Zhang, and J. Sun · 2022
Cited alongside, same era.
Swformer: Sparse window transformer for 3d object detection in point clouds
P. Sun, M. Tan, W. Wang, C. Liu, F. Xia, Z. Leng, and D. Anguelov · 2022
Cited alongside, same era.
Y. Wu, R. Li, Z. Qin, X. Zhao, and X. Li · 2023
Later among the works it cites.
Sparsefusion: Fusing multi-modal sparse representations for multi-sensor 3d object detection
Y. Xie, C. Xu, M.-J. Rakotosaona, P. Rim, F. Tombari, K. Keutzer, M. Tomizuka, and W. Zhan · 2023
Later among the works it cites.
Cross modal transformer: Towards fast and robust 3d object detection
J. Yan, Y. Liu, J. Sun, F. Jia, S. Li, T. Wang, and X. Zhang · 2023
Later among the works it cites.
Llm4drive: A survey of large language models for autonomous driving
Z. Yang, X. Jia, H. Li, and J. Yan · 2023
Later among the works it cites.
Amp: Autoregressive motion prediction revisited with next token prediction for autonomous driving
X. Jia, S. Shi, Z. Chen, L. Jiang, W. Liao, T. He, and J. Yan · 2024
Closest in time.
Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving
X. Jia, Z. Yang, Q. Li, Z. Zhang, and J. Yan · 2024
Closest in time.
Dualbev: Cnn is all you need in view transformation
P. Li, W. Shen, Q. Huang, and D. Cui · 2024
Closest in time.
Think2drive: Efficient reinforcement learning by thinking with latent world model for autonomous driving (in carla-v2)
Q. Li, X. Jia, S. Wang, and J. Yan · 2024
Closest in time.
Gafusion: Adaptive fusing lidar and camera with multiple guidance for 3d object detection
X. Li, B. Fan, J. Tian, and H. Fan · 2024
Closest in time.
Fully sparse fusion for 3d object detection
Y. Li, L. Fan, Y. Liu, Z. Huang, Y. Chen, N. Wang, and Z. Zhang · 2024
Closest in time.
Activead: Planning-oriented active learning for end-to-end autonomous driving, 2024
H. Lu, X. Jia, Y. Xie, W. Liao, X. Yang, and J. Yan · 2024
Closest in time.
Graphbev: Towards robust bev feature alignment for multi-modal 3d object detection
Z. Song, L. Yang, S. Xu, L. Liu, D. Xu, C. Jia, F. Jia, and L. Wang · 2024
Closest in time.
Point transformer v3: Simpler faster stronger
X. Wu, L. Jiang, P.-S. Wang, Z. Liu, X. Liu, Y. Qiao, W. Ouyang, T. He, and H. Zhao · 2024
Closest in time.
Is-fusion: Instance-scene collaborative fusion for multimodal 3d object detection
J. Yin, J. Shen, R. Chen, W. Li, R. Yang, P. Frossard, and W. Wang · 2024
Closest in time.
J. You, X. Jia, Z. Zhang, Y. Zhu, and J. Yan · 2024
Closest in time.
J. You, X. Jia, Z. Zhang, Y. Zhu, and J. Yan · 2024
Closest in time.
Safdnet: A simple and effective network for fully sparse 3d object detection
G. Zhang, J. Chen, G. Gao, J. Li, S. Liu, and X. Hu · 2024
Closest in time.
Hednet: A hierarchical encoder-decoder network for 3d object detection in point clouds
G. Zhang, C. Junnan, G. Gao, J. Li, and X. Hu · 2024
Closest in time.
Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions
C. Fan, X. Jia, Y. Sun, Y. Wang, J. Wei, Z. Gong, X. Zhao, M. Tomizuka, X. Yang, J. Yan, et al · 2025
Closest in time.
Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving
X. Jia, Z. Yang, Q. Li, Z. Zhang, and J. Yan · 2025
Closest in time.
Drivetransformer: Unified transformer for scalable end-to-end autonomous driving
X. Jia, J. You, Z. Zhang, and J. Yan · 2025
Closest in time.
Resim: Reliable world simulation for autonomous driving
J. Yang, K. Chitta, S. Gao, L. Chen, Y. Shao, X. Jia, H. Li, A. Geiger, X. Yue, and L. Chen · 2025
Closest in time.
Trajectory-llm: A language-based data generator for trajectory prediction in autonomous driving
K. Yang, Z. Guo, G. Lin, H. Dong, Z. Huang, Y. Wu, D. Zuo, J. Peng, Z. Zhong, X. Wang, et al · 2025
Closest in time.
Drivemoe: Mixture-of-experts for vision-language-action model in end-to-end autonomous driving
Z. Yang, Y. Chai, X. Jia, Q. Li, Y. Shao, X. Zhu, H. Su, and J. Yan · 2025
Closest in time.
Z. Yang, X. Jia, Q. Li, X. Yang, M. Yao, and J. Yan · 2025
Closest in time.
Pointobb-v3: Expanding performance boundaries of single point-supervised oriented object detection
P. Zhang, J. Luo, X. Yang, Y. Yu, Q. Li, Y. Zhou, X. Jia, X. Lu, J. Chen, X. Li, et al · 2025
Closest in time.