Fetching the paper…
Reading the bibliography…
Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Kiros, R.; Salakhutdinov, R.; and Zemel, R. S. 2014 · 2014
Earlier work this paper cites.
24/7 place recognition by view synthesis
Torii, A.; Arandjelovic, R.; Sivic, J.; Okutomi, M.; and Pajdla, T. 2015 · 2015
Earlier work this paper cites.
NetVLAD: CNN architecture for weakly supervised place recognition
Arandjelovic, R.; Gronat, P.; Torii, A.; Pajdla, T.; and Sivic, J. 2016 · 2016
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Efficient & effective prioritized matching for large-scale image-based localization
Sattler, T.; Leibe, B.; and Kobbelt, L. 2016 · 2016
Earlier work this paper cites.
Dsac-differentiable ransac for camera localization
Brachmann, E.; Krull, A.; Nowozin, S.; Shotton, J.; Michel, F.; Gumhold, S.; and Rother, C. 2017 · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Qi, C. R.; Yi, L.; Su, H.; and Guibas, L. J. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Relation Networks for Object Detection
Hu, H.; Gu, J.; Zhang, Z.; Dai, J.; and Wei, Y. 2018 · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2018 · 2018
Earlier work this paper cites.
Entangled Transformer for Image Captioning
Li, G.; Zhu, L.; Liu, P.; and Yang, Y. 2019 · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
From coarse to fine: Robust hierarchical localization at large scale
Sarlin, P.-E.; Cadena, C.; Siegwart, R.; and Dymczyk, M. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
Tan, H.; and Bansal, M. 2019 · 2019
Cited alongside, same era.
Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Achlioptas, P.; Abdelreheem, A.; Xia, F.; Elhoseiny, M.; and Guibas, L. 2020 · 2020
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; and Zagoruyko, S. 2020 · 2020
Cited alongside, same era.
Scanrefer: 3d object localization in rgb-d scans using natural language
Free-form description guided 3d visual graph network for object grounding in point cloud
Feng, M.; Li, Z.; Li, Q.; Zhang, L.; Zhang, X.; Zhu, G.; Zhang, H.; Wang, Y.; and Mian, A. 2021 · 2021
Later among the works it cites.
Fast Convergence of DETR With Spatially Modulated Co-Attention
Gao, P.; Zheng, M.; Wang, X.; Dai, J.; and Li, H. 2021 · 2021
Later among the works it cites.
Patch-netvlad: Multi-scale fusion of locally-global descriptors for place recognition
Hausler, S.; Garg, S.; Xu, M.; Milford, M.; and Fischer, T. 2021 · 2021
Later among the works it cites.
KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3D
Liao, Y.; Xie, J.; and Geiger, A. 2021 · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, D. Z.; Chang, A. X.; and Nießner, M. 2020 · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2020
Cited alongside, same era.
Semantic Correspondence as an Optimal Transport Problem
Liu, Y.; Zhu, L.; Yamada, M.; and Yang, Y. 2020 · 2020
Cited alongside, same era.
Superglue: Learning feature matching with graph neural networks
Sarlin, P.-E.; DeTone, D.; Malisiewicz, T.; and Rabinovich, A. 2020 · 2020
Cited alongside, same era.
To learn or not to learn: Visual localization from essential matrices
Zhou, Q.; Sattler, T.; Pollefeys, M.; and Leal-Taixe, L. 2020 · 2020
Cited alongside, same era.
Deformable DETR: Deformable Transformers for End-to-End Object Detection
Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; and Dai, J. 2020 · 2020
Cited alongside, same era.
Per-pixel classification is not all you need for semantic segmentation
Cheng, B.; Schwing, A.; and Kirillov, A. 2021 · 2021
Cited alongside, same era.
Rethinking and improving relative position encoding for vision transformer
Wu, K.; Peng, H.; Chen, M.; Fu, J.; and Chao, H. 2021 · 2021
Later among the works it cites.
Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Yuan, Z.; Yan, X.; Liao, Y.; Zhang, R.; Wang, S.; Li, Z.; and Cui, S. 2021 · 2021
Later among the works it cites.
Video Corpus Moment Retrieval with Contrastive Learning
Zhang, H.; Sun, A.; Jing, W.; Nan, G.; Zhen, L.; Zhou, J. T.; and Goh, R. S. M. 2021 · 2021
Later among the works it cites.
Point spatio-temporal transformer networks for point cloud video modeling
Fan, H.; Yang, Y.; and Kankanhalli, M. 2022 · 2022
Later among the works it cites.
Text2Pos: Text-to-Point-Cloud Cross-Modal Localization
Kolmet, M.; Zhou, Q.; Osep, A.; and Leal-Taixe, L. 2022 · 2022
Later among the works it cites.
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Later among the works it cites.
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; and Zhang, L. 2022 · 2022
Later among the works it cites.