Fetching the paper…
Reading the bibliography…
Vision-centric joint perception and prediction (PnP) has become an emerging trend in autonomous driving research.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Gate-variants of gated recurrent unit (gru) neural networks
Rahul Dey and Fathi M Salem · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Intentnet: Learning to predict intention from raw sensor data
Sergio Casas, Wenjie Luo, and Raquel Urtasun · 2018
Earlier work this paper cites.
Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net
Wenjie Luo, Bin Yang, and Raquel Urtasun · 2018
Earlier work this paper cites.
Monocular semantic occupancy grid mapping with convolutional variational encoder–decoder networks
Chenyang Lu, Marinus Jacobus Gerardus van de Molengraft, and Gijs Dubbelman · 2019
Earlier work this paper cites.
Scaling autoregressive video models
Dirk Weissenborn, Oscar Täckström, and Jakob Uszkoreit · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Spagnn: Spatially-aware graph neural networks for relational behavior forecasting from sensor data
Sergio Casas, Cole Gulino, Renjie Liao, and Raquel Urtasun · 2020
Earlier work this paper cites.
3d point cloud processing and learning for autonomous driving: Impacting map creation, localization, and perception
Siheng Chen, Baoan Liu, Chen Feng, Carlos Vallespi-Gonzalez, and Carl Wellington · 2020
Earlier work this paper cites.
Any motion detector: Learning class-agnostic scene dynamics from a sequence of lidar point clouds
Artem Filatov, Andrey Rykov, and Viacheslav Murashkin · 2020
Earlier work this paper cites.
Collaborative motion predication via neural motion message passing
Yue Hu, Siheng Chen, Ya Zhang, and Xiao Gu · 2020
Earlier work this paper cites.
Pillarflow: End-to-end birds-eye-view flow estimation for autonomous driving
Kuan-Hui Lee, Matthew Kliemann, Adrien Gaidon, Jie Li, Chao Fang, Sudeep Pillai, and Wolfram Burgard · 2020
Earlier work this paper cites.
Pnpnet: End-to-end perception and prediction with tracking in the loop
Ming Liang, Bin Yang, Wenyuan Zeng, Yun Chen, Rui Hu, Sergio Casas, and Raquel Urtasun · 2020
Earlier work this paper cites.
Cross-view semantic segmentation for sensing surroundings
Bowen Pan, Jiankai Sun, Ho Yin Tiga Leung, Alex Andonian, and Bolei Zhou · 2020
Earlier work this paper cites.
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d
Jonah Philion and Sanja Fidler · 2020
Earlier work this paper cites.
Ruslan Rakhimov, Denis Volkhonskiy, Alexey Artemov, Denis Zorin, and Evgeny Burnaev · 2020
Cited alongside, same era.
Predicting semantic map representations from images using pyramid occupancy networks
Thomas Roddick and Roberto Cipolla · 2020
Cited alongside, same era.
Perceive, predict, and plan: Safe motion planning through interpretable semantic representations
Abbas Sadat, Sergio Casas, Mengye Ren, Xinyu Wu, Pranaab Dhawan, and Raquel Urtasun · 2020
Cited alongside, same era.
Liranet: End-to-end trajectory prediction using spatio-temporal radar fusion
Meet Shah, Zhiling Huang, Ankit Laddha, Matthew Langford, Blake Barber, Sidney Zhang, Carlos Vallespi-Gonzalez, and Raquel Urtasun · 2020
Cited alongside, same era.
Identifying unknown instances for autonomous driving
Kelvin Wong, Shenlong Wang, Mengye Ren, Ming Liang, and Raquel Urtasun · 2020
Persformer: 3d lane detection via perspective transformer and the openlane benchmark
Li Chen, Chonghao Sima, Yang Li, Zehan Zheng, Jiajie Xu, Xiangwei Geng, Hongyang Li, Conghui He, Jianping Shi, Yu Qiao, et al · 2022
Later among the works it cites.
Rstt: Real-time spatial temporal transformer for space-time video super-resolution
Zhicheng Geng, Luming Liang, Tianyu Ding, and Ilya Zharkov · 2022
Later among the works it cites.
Maskvit: Masked visual pre-training for video prediction
Agrim Gupta, Stephen Tian, Yunzhi Zhang, Jiajun Wu, Roberto Martín-Martín, and Li Fei-Fei · 2022
Later among the works it cites.
A simple baseline for bev perception without lidar
Adam W Harley, Zhaoyuan Fang, Jie Li, Rares Ambrus, and Katerina Fragkiadaki · 2022
Later among the works it cites.
St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning
Shengchao Hu, Li Chen, Penghao Wu, Hongyang Li, Junchi Yan, and Dacheng Tao · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Motionnet: Joint perception and motion prediction for autonomous driving based on bird’s eye view maps
Pengxiang Wu, Siheng Chen, and Dimitris N Metaxas · 2020
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Cited alongside, same era.
Mp3: A unified model to map, perceive, predict and plan
Sergio Casas, Abbas Sadat, and Raquel Urtasun · 2021
Cited alongside, same era.
Latent variable sequential set transformers for joint multi-agent motion prediction
Roger Girgis, Florian Golemo, Felipe Codevilla, Martin Weiss, Jim Aldon D’Souza, Samira Ebrahimi Kahou, Felix Heide, and Christopher Pal · 2021
Cited alongside, same era.
Fiery: Future instance prediction in bird’s-eye view from surround monocular cameras
Anthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas, Jeffrey Hawke, Vijay Badrinarayanan, Roberto Cipolla, and Alex Kendall · 2021
Cited alongside, same era.
Bevdet: High-performance multi-camera 3d object detection in bird-eye-view
Junjie Huang, Guan Huang, Zheng Zhu, and Dalong Du · 2021
Cited alongside, same era.
Hdmapnet: An online hd map construction and evaluation framework, 2021
Qi Li, Yue Wang, Yilun Wang, and Hang Zhao · 2021
Cited alongside, same era.
Later among the works it cites.
Where2comm: Communication-efficient collaborative perception via spatial confidence maps
Yue Hu, Shaoheng Fang, Zixing Lei, Yiqi Zhong, and Siheng Chen · 2022
Later among the works it cites.
Time3d: End-to-end joint monocular 3d object detection and tracking for autonomous driving
Peixuan Li and Jieyu Jin · 2022
Later among the works it cites.
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, and Jifeng Dai · 2022
Later among the works it cites.
Bevfusion: A simple and robust lidar-camera fusion framework
Tingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia, Zhiwei Lin, Yongtao Wang, Tao Tang, Bing Wang, and Zhi Tang · 2022
Later among the works it cites.
Petrv2: A unified framework for 3d perception from multi-camera images
Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Qi Gao, Tiancai Wang, Xiangyu Zhang, and Jian Sun · 2022
Later among the works it cites.
Video frame interpolation with transformer
Liying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu, and Jiaya Jia · 2022
Later among the works it cites.
Bevsegformer: Bird’s eye view semantic segmentation from arbitrary camera rigs
Lang Peng, Zhirong Chen, Zhangjie Fu, Pengpeng Liang, and Erkang Cheng · 2022
Later among the works it cites.
Zequn Qin, Jingyu Chen, Chao Chen, Xiaozhi Chen, and Xi Li · 2022
Later among the works it cites.
Translating images into maps
Avishkar Saha, Oscar Mendez, Chris Russell, and Richard Bowden · 2022
Later among the works it cites.
Video frame interpolation transformer
Zhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen, and Ming-Hsuan Yang · 2022
Later among the works it cites.
Be-sti: Spatial-temporal integrated network for class-agnostic motion prediction with bidirectional enhancement
Yunlong Wang, Hongyu Pan, Jun Zhu, Yu-Huan Wu, Xin Zhan, Kun Jiang, and Diange Yang · 2022
Later among the works it cites.
Groupnet: Multiscale hypergraph neural networks for trajectory prediction with relational reasoning
Chenxin Xu, Maosen Li, Zhenyang Ni, Ya Zhang, and Siheng Chen · 2022
Later among the works it cites.
Beverse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving
Yunpeng Zhang, Zheng Zhu, Wenzhao Zheng, Junjie Huang, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Later among the works it cites.
Cross-view transformers for real-time map-view semantic segmentation
Brady Zhou and Philipp Krähenbühl · 2022
Later among the works it cites.