Fetching the paper…
Reading the bibliography…
In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
3d object proposals for accurate object class detection
Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, and et. al · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric A Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Multi-view 3d object detection network for autonomous driving
Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz · 2018
Earlier work this paper cites.
Stereo r-cnn based 3d object detection for autonomous driving
Peiliang Li, Xiaozhi Chen, and Shaojie Shen · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving
Yan Wang, Wei-Lun Chao, Divyansh Garg, and et. al · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, and et.al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Collaborative unsupervised domain adaptation for medical image diagnosis
Yifan Zhang, Ying Wei, Qingyao Wu, Peilin Zhao, Shuaicheng Niu, Junzhou Huang, and Mingkui Tan · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Earlier work this paper cites.
Magicdrive: 3d controllable driving video synthesis using layout and 3d representations
Yuming Hong, Weiyao Xu, Yi Yang, Hao Zhou, and Xiaoxue Wang · 2021
Earlier work this paper cites.
Drivingdiffusion: Realistic multi-view video generation with diffusion models
Junsoo Kim, Yeongjae Ye, Dongheon Cho, Jaechan Seo, and Jiseob Jeong · 2021
Cited alongside, same era.
Vpd: Vision-based perception diffusion models for human pose estimation and dense visual annotations
Feng Li, Yuan Zhang, Xinyang Wang, and Meng Liu · 2021
Cited alongside, same era.
Source-free domain adaptation via avatar prototype generation and adaptation
Zhen Qiu, Yifan Zhang, Hongbin Lin, Shuaicheng Niu, Yanxia Liu, Qing Du, and Mingkui Tan · 2021
Cited alongside, same era.
Videogpt: Video generation using vq-vae and transformers
Xu Yan, Weiming Zeng, Jiaming Zhu, Hailin Chen, Yaowei Li, and Deng Cai · 2021
Cited alongside, same era.
Center-based 3d object detection and tracking
Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl · 2021
Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation
Zhijian Liu, Haotian Tang, Alexander Amini, Xinyu Yang, Huizi Mao, Daniela L Rus, and Song Han · 2023
Later among the works it cites.
Unicontrol: A unified diffusion model for controllable visual generation in the wild
Yuheng Qin, Zhaoyang Zhang, Han Zhang, Dong Li, Baining Chen, Lu Zhang, Jiuxiang Gu, and Dongdong Zhang · 2023
Later among the works it cites.
Plug-and-play diffusion features for text-driven image-to-image translation
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel · 2023
Later among the works it cites.
Virtual sparse convolution for multimodal 3d object detection
Hai Wu, Chenglu Wen, Shaoshuai Shi, Xin Li, and Cheng Wang · 2023
Later among the works it cites.
Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Objects are different: Flexible monocular 3d object detection
Yunpeng Zhang, Jiwen Lu, and Jie Zhou · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Diffumask: Synthetic image and semantic mask annotation using cross-attention in diffusion models
Xinyu Huang, Mingyu Chen, Yunqi Li, Hao Wang, and Yuxing Zhao · 2022
Cited alongside, same era.
Prototype-guided continual adaptation for class-incremental unsupervised domain adaptation
Hongbin Lin, Yifan Zhang, Zhen Qiu, Shuaicheng Niu, Chuang Gan, Yanxia Liu, and Mingkui Tan · 2022
Cited alongside, same era.
Multi-modal fusion for 3d object detection in autonomous driving
Xiaoyang Liu, Yikang Wang, Xinyu Zhang, Wei Ma, Chao Wang, and Yun Chen · 2022
Cited alongside, same era.
Temporal fusion with attention mechanisms for driving scenarios
Seokhwa Moon, Sanghyuk Hwang, Hyunsu Lee, and Minje Kim · 2022
Cited alongside, same era.
Monoground: Detecting monocular 3d objects from the ground
Zequn Qin and Xi Li · 2022
Cited alongside, same era.
Mononerd: Nerf-like representations for monocular 3d object detection
Junkai Xu, Liang Peng, Haoran Cheng, Hao Li, Wei Qian, Ke Li, Wenxiao Wang, and Deng Cai · 2023
Later among the works it cites.
Hipa: enabling one-step text-to-image diffusion models via high-frequency-promoting adaptation
Yifan Zhang and Bryan Hooi · 2023
Later among the works it cites.
Uni-controlnet: All-in-one control to text-to-image diffusion models
Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Dongdong Zhang, Lu Yuan, Han Zhang, Dong Li, Baining Chen, and Lu Zhang · 2023
Later among the works it cites.
Layoutdiffusion: Controllable diffusion model for layout-to-image generation
Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li · 2023
Later among the works it cites.
Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition
Sicheng Mo, Fangzhou Mu, Kuan Heng Lin, Yanli Liu, Bochen Guan, Yin Li, and Bolei Zhou · 2024
Later among the works it cites.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al · 2024
Later among the works it cites.
Robofusion: towards robust multi-modal 3d object detection via sam
Ziying Song, Guoxing Zhang, Lin Liu, Lei Yang, Shaoqing Xu, Caiyan Jia, Feiyang Jia, and Li Wang · 2024
Later among the works it cites.
Street-view image generation from a bird’s-eye view layout
Alexander Swerdlow, Runsheng Xu, and Bolei Zhou · 2024
Later among the works it cites.
Monocd: Monocular 3d object detection with complementary depths
Longfei Yan, Pei Yan, Shengzhou Xiong, Xuanyu Xiang, and Yihua Tan · 2024
Later among the works it cites.
Memo: Memory-guided diffusion for expressive talking video generation
Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan, Zhenxiong Tan, Jiahao Lu, Chuanxin Tang, Bo An, and Shuicheng Yan · 2024
Later among the works it cites.
Monotta: Fully test-time adaptation for monocular 3d object detection
Hongbin Lin, Yifan Zhang, Shuaicheng Niu, Shuguang Cui, and Zhen Li · 2025
Closest in time.
Monowad: Weather-adaptive diffusion model for robust monocular 3d object detection
Youngmin Oh, Hyung-Il Kim, Seong Tae Kim, and Jung Uk Kim · 2025
Closest in time.
Benchmarking and improving bird’s eye view perception robustness in autonomous driving
Shaoyuan Xie, Lingdong Kong, Wenwei Zhang, Jiawei Ren, Liang Pan, Kai Chen, and Ziwei Liu · 2025
Closest in time.