Fetching the paper…
Reading the bibliography…
We propose an end-to-end network that takes a single perspective RGB image of a complex road scene as input, to produce occlusion-reasoned layouts in perspective space as well as a parametric bird's-eye-view (BEV) space.
Multiple view geometry in computer vision
Richard Hartley and Andrew Zisserman · 2003
Earlier work this paper cites.
Decomposing a scene into geometric and semantically consistent regions
Stephen Gould, Richard Fulton, and Daphne Koller · 2009
Earlier work this paper cites.
Single image depth estimation from predicted semantic labels
Beyang Liu, Stephen Gould, and Daphne Koller · 2010
Earlier work this paper cites.
Beyond the line of sight: labeling the underlying surfaces
Ruiqi Guo and Derek Hoiem · 2012
Earlier work this paper cites.
Automatic Dense Visual Semantic Mapping from Street-Level Imagery
Sunando Sengupta, Paul Sturgess, L̀ubor Ladický, and Philip H. S. Torr · 2012
Earlier work this paper cites.
Vision meets Robotics: The KITTI Dataset
Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun · 2013
Earlier work this paper cites.
Semantic image segmentation with deep convolutional nets and fully connected crfs
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille · 2014
Earlier work this paper cites.
3D Traffic Scene Understanding from Movable Platforms
Andreas Geiger, Martin Lauer, Christian Wojek, Christoph Stiller, and Raquel Urtasun · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Scene Parsing with Object Instances and Occlusion Ordering
Joseph Tighe, Marc Niethammer, and Svetlana Lazebnik · 2014
Earlier work this paper cites.
Deeply-Supervised Nets
Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu · 2015
Earlier work this paper cites.
Multiclass semantic video segmentation with object-level active inference
Buyu Liu and Xuming He · 2015
Earlier work this paper cites.
Rent3D: Floor-Plan Priors for Monocular Layout Estimation
Chenxi Liu, Alexander G. Schwing, Kaustav Kundu, Raquel Urtasun, and Sanja Fidler · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
3D Semantic Parsing of Large-Scale Indoor Spaces
Iro Armeni, Ozan Sener, Amir R. Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese · 2016
Earlier work this paper cites.
A Continuous Occlusion Model for Road Scene Understanding
Vikas Dhiman, Quoc-Huy Tran, Jason J. Corso, and Manmohan Chandraker · 2016
Earlier work this paper cites.
Unsupervised cnn for single view depth estimation: Geometry to the rescue
Ravi Garg, Vijay Kumar Bg, Gustavo Carneiro, and Ian Reid · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Learning from Maps: Visual Common Sense for Autonomous Driving
Ari Seff and Jianxiong Xiao · 2016
Cited alongside, same era.
Spatiotemporal multiplier networks for video action recognition
Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes · 2017
Cited alongside, same era.
Unsupervised monocular depth estimation with left-right consistency
Clément Godard, Oisin Mac Aodha, and Gabriel J Brostow · 2017
Cited alongside, same era.
Cognitive Mapping and Planning for Visual Navigation
Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik · 2017
Stereo r-cnn based 3d object detection for autonomous driving
Peiliang Li, Xiaozhi Chen, and Shaojie Shen · 2019
Later among the works it cites.
Monocular Semantic Occupancy Grid Mapping With Convolutional Variational Encoder-Decoder Networks
Chenyang Lu, Marinus Jacobus Gerardus van de Molengraft, and Gijs Dubbelman · 2019
Later among the works it cites.
Spatial-aware feature aggregation for image based cross-view geo-localization
Yujiao Shi, Liu Liu, Xin Yu, and Hongdong Li · 2019
Later among the works it cites.
Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving
Yan Wang, Wei-Lun Chao, Divyansh Garg, Bharath Hariharan, Mark Campbell, and Kilian Q Weinberger · 2019
Later among the works it cites.
A dataset for high-level 3d scene understanding of complex road scenes in the top-view
Ziyan Wang, Buyu Liu, Samuel Schulter, and Manmohan Chandraker · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep supervision with shape concepts for occlusion-aware 3d object parsing
Chi Li, M. Zeeshan Zia, Quoc-Huy Tran, Xiang Yu, Gregory D. Hager, and Manmohan Chandraker · 2017
Cited alongside, same era.
Flow-guided feature aggregation for video object detection
Xizhou Zhu, Yujie Wang, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Cited alongside, same era.
Deep feature flow for video recognition
Xizhou Zhu, Yuwen Xiong, Jifeng Dai, Lu Yuan, and Yichen Wei · 2017
Cited alongside, same era.
Reading between the Lanes: Road Layout Reconstruction from Partially Segmented Scenes
Lars Kunze, Tom Bruls, Tarlan Suleymanov, and Paul Newman · 2018
Cited alongside, same era.
The NuScenes data set
NuTonomy · 2018
Cited alongside, same era.
Orthographic feature transform for monocular 3d object detection
Thomas Roddick, Alex Kendall, and Roberto Cipolla · 2018
Cited alongside, same era.
A parametric top-view representation of complex road scenes
Ziyan Wang, Buyu Liu, Samuel Schulter, and Manmohan Chandraker · 2019
Later among the works it cites.
Understanding road layout from videos as a whole
Buyu Liu, Bingbing Zhuang, Samuel Schulter, Pan Ji, and Manmohan Chandraker · 2020
Later among the works it cites.
Monolayout: Amodal scene layout from a single image
Kaustubh Mani, Swapnil Daga, Shubhika Garg, Sai Shankar Narasimhan, Madhava Krishna, and Krishna Murthy Jatavallabhula · 2020
Later among the works it cites.
Autolay: Benchmarking amodal layout estimation for autonomous driving
Kaustubh Mani, N Sai Shankar, Krishna Murthy Jatavallabhula, and K Madhava Krishna · 2020
Later among the works it cites.
Cross-view semantic segmentation for sensing surroundings
B. Pan, J. Sun, H. Y. T. Leung, A. Andonian, and B. Zhou · 2020
Later among the works it cites.
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d
Jonah Philion and Sanja Fidler · 2020
Later among the works it cites.
Predicting semantic map representations from images using pyramid occupancy networks
Thomas Roddick and Roberto Cipolla · 2020
Later among the works it cites.
Deep high-resolution representation learning for visual recognition
Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al · 2020
Later among the works it cites.
Self-supervised scene de-occlusion
Xiaohang Zhan, Xingang Pan, Bo Dai, Ziwei Liu, Dahua Lin, and Chen Change Loy · 2020
Later among the works it cites.
Structured bird’s-eye-view traffic scene understanding from on board images
Yigit Baran Can, Alexander Liniger, Danda Pani Paudel, and Luc Van Gool · 2021
Closest in time.
Projecting your view attentively: Monocular road scene layout estimation via cross-view transformation
Weixiang Yang, Qi Li, Wenxi Liu, Yuanlong Yu, Yuexin Ma, Shengfeng He, and Jia Pan · 2021
Closest in time.