Fetching the paper…
Reading the bibliography…
A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun · 2013
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Current challenges in autonomous driving
I. Barabas, A. Todoruţ, N. Cordoş, and A. Molea · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi, H. Su, K. Mo, and L. J. Guibas · 2017
Earlier work this paper cites.
Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst
M. Bansal, A. Krizhevsky, and A. Ogale · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
J. Kim, A. Rohrbach, T. Darrell, J. Canny, and Z. Akata · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2018
Earlier work this paper cites.
Pointpillars: Fast encoders for object detection from point clouds
A. H. Lang, S. Vora, H. Caesar, L. Zhou, J. Yang, and O. Beijbom · 2019
Earlier work this paper cites.
End-to-end interpretable neural motion planner
W. Zeng, W. Luo, S. Suo, A. Sadat, B. Yang, S. Casas, and R. Urtasun · 2019
Earlier work this paper cites.
Talk2car: Taking control of your self-driving car
T. Deruyttere, S. Vandenhende, D. Grujicic, L. Van Gool, and M.-F. Moens · 2019
Earlier work this paper cites.
Vectornet: Encoding hd maps and agent dynamics from vectorized representation
J. Gao, C. Sun, H. Zhao, Y. Shen, D. Anguelov, C. Li, and C. Schmid · 2020
Earlier work this paper cites.
End-to-end model-free reinforcement learning for urban driving using implicit affordances
M. Toromanoff, E. Wirbel, and F. Moutarde · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom · 2020
Earlier work this paper cites.
Explainable object-induced action decision for autonomous vehicles
Y. Xu, X. Yang, L. Gong, H.-C. Lin, T.-Y. Wu, Y. Li, and N. Vasconcelos · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
Tnt: Target-driven trajectory prediction
H. Zhao, J. Gao, T. Lan, C. Sun, B. Sapp, B. Varadarajan, Y. Shen, Y. Shen, Y. Chai, C. Schmid, C. Li, and D. Anguelov · 2021
Earlier work this paper cites.
Multimodal motion prediction with stacked transformers
Y. Liu, J. Zhang, L. Fang, Q. Jiang, and B. Zhou · 2021
Earlier work this paper cites.
Densetnt: End-to-end trajectory prediction from dense goal sets
J. Gu, C. Sun, and H. Zhao · 2021
Earlier work this paper cites.
Learning to drive from a world on rails
D. Chen, V. Koltun, and P. Krähenbühl · 2021
Earlier work this paper cites.
Perceive, attend, and drive: Learning spatial attention for safe self-driving
B. Wei, M. Ren, W. Zeng, M. Liang, B. Yang, and R. Urtasun · 2021
Earlier work this paper cites.
Safe local motion planning with self-supervised freespace forecasting
P. Hu, A. Huang, J. Dolan, D. Held, and D. Ramanan · 2021
Earlier work this paper cites.
Mp3: A unified model to map, perceive, predict and plan
S. Casas, A. Sadat, and R. Urtasun · 2021
Earlier work this paper cites.
Sutd-trafficqa: A question answering benchmark and an efficient network for video reasoning over traffic events
L. Xu, H. Huang, and J. Liu · 2021
Earlier work this paper cites.
Detr3d: 3d object detection from multi-view images via 3d-to-2d queries
Y. Wang, V. C. Guizilini, T. Zhang, Y. Wang, H. Zhao, and J. Solomon · 2022
Earlier work this paper cites.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Z. Li, W. Wang, H. Li, E. Xie, C. Sima, T. Lu, Y. Qiao, and J. Dai · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Plant: Explainable planning transformers via object-level representations
K. Renz, K. Chitta, O.-B. Mercea, A. Koepke, Z. Akata, and A. Geiger · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Differentiable raycasting for self-supervised occupancy forecasting
T. Khurana, P. Hu, A. Dave, J. Ziglar, D. Held, and D. Ramanan · 2022
Gaia-1: A generative world model for autonomous driving
A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado · 2023
Later among the works it cites.
A survey of large language models for autonomous driving
Z. Yang, X. Jia, H. Li, and J. Yan · 2023
Later among the works it cites.
Referring multi-object tracking
D. Wu, W. Han, T. Wang, X. Dong, X. Zhang, and J. Shen · 2023
Later among the works it cites.
Language prompt for autonomous driving
D. Wu, W. Han, T. Wang, Y. Liu, X. Zhang, and J. Shen · 2023
Later among the works it cites.
Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario
T. Qian, J. Chen, L. Zhuo, Y. Jiao, and Y.-G. Jiang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning
S. Hu, L. Chen, P. Wu, H. Li, J. Yan, and D. Tao · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Cited alongside, same era.
Wayformer: Motion forecasting via simple & efficient attention networks
N. Nayakanti, R. Al-Rfou, A. Zhou, K. Goel, K. S. Refaat, and B. Sapp · 2023
Cited alongside, same era.
Uncertainty-aware decision transformer for stochastic driving environments
Z. Li, F. Nie, Q. Sun, F. Da, and H. Zhao · 2023
Cited alongside, same era.
What matters in training a gpt4-style language model with multimodal inputs?
Y. Zeng, H. Zhang, J. Zheng, J. Xia, G. Wei, Y. Wei, Y. Zhang, and T. Kong · 2023
Cited alongside, same era.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Z. Xu, Y. Zhang, E. Xie, Z. Zhao, Y. Guo, K. K. Wong, Z. Li, and H. Zhao · 2023
Cited alongside, same era.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Cited alongside, same era.
Later among the works it cites.
Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning
E. Sachdeva, N. Agarwal, S. Chundi, S. Roelofs, J. Li, B. Dariush, C. Choi, and M. Kochenderfer · 2023
Later among the works it cites.
Drama: Joint risk localization and captioning in driving
S. Malla, C. Choi, I. Dwivedi, J. H. Choi, and J. Li · 2023
Later among the works it cites.
Vad: Vectorized scene representation for efficient autonomous driving
B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes
J.-T. Zhai, Z. Feng, J. Du, Y. Mao, J.-J. Liu, Z. Tan, Y. Zhang, X. Ye, and J. Wang · 2023
Later among the works it cites.
Seed-bench: Benchmarking multimodal llms with generative comprehension
B. Li, R. Wang, G. Wang, Y. Ge, Y. Ge, and Y. Shan · 2023
Later among the works it cites.
Drivelm: Driving with graph visual question answering
C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, P. Luo, A. Geiger, and H. Li · 2023
Later among the works it cites.
Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices
X. Chu, L. Qiao, X. Lin, S. Xu, Y. Yang, Y. Hu, F. Wei, X. Zhang, B. Zhang, X. Wei, et al · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Medusa: Simple framework for accelerating llm generation with multiple decoding heads, 2023
T. Cai, Y. Li, Z. Geng, H. Peng, and T. Dao · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Closest in time.
Grok-1.5 vision preview
X.ai · 2024
Closest in time.
Introducing qwen1.5, February 2024
Q. Team · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al · 2024
Closest in time.
Phi-3 technical report: A highly capable language model locally on your phone
M. Abdin, S. A. Jacobs, A. A. Awan, J. Aneja, A. Awadallah, H. Awadalla, N. Bach, A. Bahree, A. Bakhtiari, H. Behl, et al · 2024
Closest in time.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2024
Closest in time.
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
B. He, H. Li, Y. K. Jang, M. Jia, X. Cao, A. Shah, A. Shrivastava, and S.-N. Lim · 2024
Closest in time.
Eagle: Speculative sampling requires rethinking feature uncertainty
Y. Li, F. Wei, C. Zhang, and H. Zhang · 2024
Closest in time.