Fetching the paper…
Reading the bibliography…
We introduce a novel visual question answering (VQA) task in the context of autonomous driving, aiming to answer natural language questions based on street-view clues.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H.; and Bansal, M. 2019 · 1908
Earlier work this paper cites.
Talk2car: Taking control of your self-driving car
Deruyttere, T.; Vandenhende, S.; Grujicic, D.; Van Gool, L.; and Moens, M.-F. 2019 · 1909
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Geiger, A.; Lenz, P.; and Urtasun, R. 2012 · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.-J.; Shamma, D.; Bernstein, M.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Lu, J.; Yang, J.; Batra, D.; and Parikh, D. 2016 · 2016
Earlier work this paper cites.
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Dai, A.; Chang, A. X.; Savva, M.; Halber, M.; Funkhouser, T.; and Nießner, M. 2017 · 2017
Earlier work this paper cites.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; and Kim, G. 2017 · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J.; Hariharan, B.; Van Der Maaten, L.; Fei-Fei, L.; Lawrence Zitnick, C.; and Girshick, R. 2017 · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Cited alongside, same era.
Uncovering the temporal context for video question answering
Zhu, L.; Xu, Z.; Yang, Y.; and Hauptmann, A. G. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Embodied question answering
Das, A.; Datta, S.; Gkioxari, G.; Lee, S.; Parikh, D.; and Batra, D. 2018 · 2018
Cited alongside, same era.
3d semantic segmentation with submanifold sparse convolutional networks
Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering
Jiang, J.; Chen, Z.; Lin, H.; Zhao, X.; and Gao, Y. 2020 · 2020
Later among the works it cites.
Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d
Philion, J.; and Fidler, S. 2020 · 2020
Later among the works it cites.
Bevdet: High-performance multi-camera 3d object detection in bird-eye-view
Huang, J.; Huang, G.; Zhu, Z.; Ye, Y.; and Du, D. 2021 · 2021
Later among the works it cites.
Yan, X.; Yuan, Z.; Du, Y.; Liao, Y.; Guo, Y.; Li, Z.; and Cui, S. 2021 · 2021
Later among the works it cites.
Center-based 3d object detection and tracking
Yin, T.; Zhou, X.; and Krahenbuhl, P. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Graham, B.; Engelcke, M.; and Van Der Maaten, L. 2018 · 2018
Cited alongside, same era.
Bilinear attention networks
Kim, J.-H.; Jun, J.; and Zhang, B.-T. 2018 · 2018
Cited alongside, same era.
Tvqa: Localized, compositional video question answering
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. L. 2018 · 2018
Cited alongside, same era.
Learning to count objects in natural images for visual question answering
Zhang, Y.; Hare, J.; and Prügel-Bennett, A. 2018 · 2018
Cited alongside, same era.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Zhou, Y.; and Tuzel, O. 2018 · 2018
Cited alongside, same era.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Cited alongside, same era.
Deep modular co-attention networks for visual question answering
Yu, Z.; Yu, J.; Cui, Y.; Tao, D.; and Tian, Q. 2019 · 2019
Cited alongside, same era.
Vinvl: Revisiting visual representations in vision-language models
Zhang, P.; Li, X.; Hu, X.; Yang, J.; Zhang, L.; Wang, L.; Choi, Y.; and Gao, J. 2021 · 2021
Later among the works it cites.
ScanQA: 3D question answering for spatial scene understanding
Azuma, D.; Miyanishi, T.; Kurita, S.; and Kawanabe, M. 2022 · 2022
Later among the works it cites.
3D question answering
Ye, S.; Chen, D.; Han, S.; and Liao, J. 2022 · 2022
Later among the works it cites.
Towards Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong Baseline
Zhao, L.; Cai, D.; Zhang, J.; Sheng, L.; Xu, D.; Zheng, R.; Zhao, Y.; Wang, L.; and Fan, X. 2022 · 2022
Later among the works it cites.
VoxelNeXt: Fully Sparse VoxelNet for 3D Object Detection and Tracking
Chen, Y.; Liu, J.; Zhang, X.; Qi, X.; and Jia, J. 2023 · 2023
Closest in time.
Referring Multi-Object Tracking
Dongming, W.; Wencheng, H.; Tiancai, W.; Xingping, D.; Xiangyu, Z.; and Shen, J. 2023 · 2023
Closest in time.
BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird’s-Eye View Representation
Liu, Z.; Tang, H.; Amini, A.; Yang, X.; Mao, H.; Rus, D.; and Han, S. 2023 · 2023
Closest in time.
SQA3D: Situated Question Answering in 3D Scenes
Ma, X.; Yong, S.; Zheng, Z.; Li, Q.; Liang, Y.; Zhu, S.-C.; and Huang, S. 2023 · 2023
Closest in time.