Fetching the paper…
Reading the bibliography…
Modern approaches for vision-centric environment perception for autonomous navigation make extensive use of self-supervised monocular depth estimation algorithms that output disparity maps.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
Vision meets robotics: the kitti dataset
Andreas Geiger, P Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Unsupervised monocular depth estimation with left-right consistency
Clement Godard, Oisin Aodha, and Gabriel Brostow · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
1 Year, 1000km: The Oxford RobotCar Dataset
Will Maddern, Geoff Pascoe, Chris Linegar, and Paul Newman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Megadepth: Learning single-view depth prediction from internet photos
Zhengqi Li and Noah Snavely · 2018
Earlier work this paper cites.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Yin Zhou and Oncel Tuzel · 2018
Earlier work this paper cites.
Semantickitti: A dataset for semantic scene understanding of lidar sequences
Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jürgen Gall · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Digging into self-supervised monocular depth estimation
Clément Godard, Oisin Aodha, Michael Firman, and Gabriel Brostow · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Does computer vision matter for action?
Brady Zhou, Philipp Krähenbühl, and Vladlen Koltun · 2019
Earlier work this paper cites.
Ford multi-av seasonal dataset
Siddharth Agarwal, Ankit Vora, Gaurav Pandey, Wayne Williams, Helen Kourous, and James McBride · 2020
Earlier work this paper cites.
The oxford radar robotcar dataset: A radar extension to the oxford robotcar dataset
Dan Barnes, Matthew Gadd, Paul Murcutt, Paul Newman, and Ingmar Posner · 2020
Earlier work this paper cites.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Cited alongside, same era.
3d sketch-aware semantic scene completion via semi-supervised structure prior
Xiaokang Chen, Kwan-Yee Lin, Chen Qian, Gang Zeng, and Hongsheng Li · 2020
Cited alongside, same era.
A2D2: Audi Autonomous Driving Dataset
Jakob Geyer, Yohannes Kassahun, Mentar Mahmudi, Xavier Ricou, Rupesh Durgesh, Andrew S. Chung, Lorenz Hauswald, Viet Hoang Pham, Maximilian Mühlegg, Sebastian Dorn, Tiffany Fernandez, Martin Jänicke, Sudesh Mirashi, Chiragkumar Savani, Martin Sturm, Oleksandr Vorobiov, Martin Oelker, Sebastian Garreis, and Peter Schuberth · 2020
Cited alongside, same era.
3d packing for self-supervised monocular depth estimation
Vitor Guizilini, Rareș Ambruș, Sudeep Pillai, Allan Raventos, and Adrien Gaidon · 2020
Cited alongside, same era.
Anisotropic convolutional networks for 3d semantic scene completion
Jie Li, Kai Han, Peng Wang, Yu Liu, and Xia Yuan · 2020
Cited alongside, same era.
One million scenes for autonomous driving: Once dataset
Jiageng Mao, Minzhe Niu, Chenhan Jiang, Hanxue Liang, Xiaodan Liang, Yamin Li, Chaoqiang Ye, Wei Zhang, Zhenguo Li, Jie Yu, et al · 2021
Later among the works it cites.
Boosting monocular depth estimation models to high-resolution via content-adaptive multi-resolution merging
S. Mahdi H. Miangoleh, Sebastian Dille, Long Mai, Sylvain Paris, and Yağız Aksoy · 2021
Later among the works it cites.
Canadian adverse driving conditions dataset
Matthew Pitropov, Danson Evan Garcia, Jason Rebello, Michael Smart, Carlos Wang, Krzysztof Czarnecki, and Steven Waslander · 2021
Later among the works it cites.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Later among the works it cites.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Later among the works it cites.
The temporal opportunist: Self-supervised multi-frame monocular depth
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amvnet: Assertion-based multi-view fusion network for lidar semantic segmentation, 2020
Venice Erin Liong, Thi Ngoc Tho Nguyen, Sergi Widjaja, Dhananjai Sharma, and Zhuang Jie Chong · 2020
Cited alongside, same era.
A*3d dataset: Towards autonomous driving in challenging environments
Quang-Hieu Pham, Pierre Sevestre, Ramanpreet Singh Pahwa, Huijing Zhan, Chun Ho Pang, Yuda Chen, Armin Mustafa, Vijay Chandrasekhar, and Jie Lin · 2020
Cited alongside, same era.
A sim2real deep learning approach for the transformation of images from multiple vehicle-mounted cameras to a semantically segmented image in bird’s eye view
Lennart Reiher, Bastian Lampe, and Lutz Eckstein · 2020
Cited alongside, same era.
Predicting semantic map representations from images using pyramid occupancy networks
Thomas Roddick and Roberto Cipolla · 2020
Cited alongside, same era.
Lmscnet: Lightweight multiscale 3d semantic completion
Luis Roldão, Raoul de Charette, and Anne Verroust-Blondet · 2020
Cited alongside, same era.
Searching efficient 3d architectures with sparse point-voxel convolution
Haotian Tang, Zhijian Liu, Shengyu Zhao, Yujun Lin, Ji Lin, Hanrui Wang, and Song Han · 2020
Cited alongside, same era.
Sparse single sweep lidar point cloud segmentation via learning contextual shape priors from scene completion
Xu Yan, Jiantao Gao, Jie Li, Ruimao Zhang, Zhen Li, Rui Huang, and Shuguang Cui · 2020
Cited alongside, same era.
J. Watson, O. Mac Aodha, V. Prisacariu, G. Brostow, and M. Firman · 2021
Later among the works it cites.
Sparse single sweep lidar point cloud segmentation via learning contextual shape priors from scene completion
Xu Yan, Jiantao Gao, Jie Li, Ruimao Zhang, Zhen Li, Rui Huang, and Shuguang Cui · 2021
Later among the works it cites.
Drinet++: Efficient voxel-as-point point cloud segmentation, 2021
Maosheng Ye, Rui Wan, Shuangjie Xu, Tongyi Cao, and Qifeng Chen · 2021
Later among the works it cites.
Learning to recover 3d scene shape from a single image
Wei Yin, Jianming Zhang, Oliver Wang, Simon Niklaus, Long Mai, Simon Chen, and Chunhua Shen · 2021
Later among the works it cites.
Point transformer
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H.S. Torr, and Vladlen Koltun · 2021
Later among the works it cites.
Cylindrical and asymmetrical 3d convolution networks for lidar segmentation
X. Zhu, H. Zhou, T. Wang, F. Hong, Y. Ma, W. Li, H. Li, and D. Lin · 2021
Later among the works it cites.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, and Jifeng Dai · 2022
Later among the works it cites.
Gcndepth: Self-supervised monocular depth estimation based on graph convolutional network
Armin Masoumian, Hatem A Rashwan, Saddam Abdulwahab, Julian Cristiano, M Salman Asif, and Domenec Puig · 2022
Later among the works it cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun · 2022
Later among the works it cites.
Lidarmultinet: Towards a unified multi-task network for lidar perception
Dongqiangzi Ye, Zixiang Zhou, Weijia Chen, Yufei Xie, Yu Wang, Panqu Wang, and Hassan Foroosh · 2022
Later among the works it cites.
Tri-perspective view for vision-based 3d semantic occupancy prediction
Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie Zhou, and Jiwen Lu · 2023
Closest in time.
Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation
Zhijian Liu, Haotian Tang, Alexander Amini, Xingyu Yang, Huizi Mao, Daniela Rus, and Song Han · 2023
Closest in time.