Fetching the paper…
Reading the bibliography…
Due to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored.
Neural machine translation by jointly learning to align and translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
SUN RGB-D: A RGB-D scene understanding benchmark suite
Song, S.; Lichtenberg, S. P.; and Xiao, J. 2015 · 2015
Earlier work this paper cites.
SceneNet: An annotated model generator for indoor scene understanding
Handa, A.; Pătrăucean, V.; Stent, S.; and Cipolla, R. 2016 · 2016
Earlier work this paper cites.
Semantic Scene Completion from a Single Depth Image
Song, S.; Yu, F.; Zeng, A.; Chang, A. X.; Savva, M.; and Funkhouser, T. 2016 · 2016
Earlier work this paper cites.
ScanNet: Richly-Annotated 3D Reconstructions of Indoor Scenes
Dai, A.; Chang, A. X.; Savva, M.; Halber, M.; Funkhouser, T.; and Nießner, M. 2017 · 2017
Earlier work this paper cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Stacked Cross Attention for Image-Text Matching
Lee, K.-H.; Chen, X.; Hua, G.; Hu, H.; and He, X. 2018 · 2018
Earlier work this paper cites.
InteriorNet: Mega-scale Multi-sensor Photo-realistic Indoor Scenes Dataset
Li, W.; Saeedi, S.; McCormac, J.; Clark, R.; Tzoumanikas, D.; Ye, Q.; Huang, Y.; Tang, R.; and Leutenegger, S. 2018 · 2018
Earlier work this paper cites.
Multi-Similarity Loss With General Pair Weighting for Deep Metric Learning
Wang, X.; Han, X.; Huang, W.; Dong, D.; and Scott, M. R. 2019 · 2019
Earlier work this paper cites.
Language-Agnostic Visual-Semantic Embeddings
Wehrmann, J.; Lopes, M. A.; Souza, D.; and Barros, R. 2019 · 2019
Cited alongside, same era.
Multi-class joint subspace learning for cross-modal retrieval
Yu, E.; Li, J.; Wang, L.; Zhang, J.; Wan, W.; and Sun, J. 2020 · 2020
Cited alongside, same era.
Similarity Reasoning and Filtration for Image-Text Matching
Diao, H.; Zhang, Y.; Ma, L.; and Lu, H. 2021 · 2021
Cited alongside, same era.
Deep Metric Learning with Self-Supervised Ranking
Fu, Z.; Li, Y.; Mao, Z.; Wang, Q.; and Zhang, Y. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Self-Supervised Synthesis Ranking for Deep Metric Learning
Fu, Z.; Mao, Z.; Yan, C.; Liu, A.-A.; Xie, H.; and Zhang, Y. 2022 · 2022
Later among the works it cites.
Intra-Class Adaptive Augmentation With Neighbor Correction for Deep Metric Learning
Fu, Z.; Mao, Z.; Hu, B.; Liu, A.-A.; and Zhang, Y. 2023 · 2023
Later among the works it cites.
Deep Hierarchical Feature Learning on Point Sets in a Metric Space
Ruizhongtai, Q. C.; Li, Y.; Hao, S.; and PointNet+, G. L. J. 2023 · 2023
Later among the works it cites.
Less is Better: Exponential Loss for Cross-Modal Matching
Wei, J.; Yang, Y.; Xu, X.; Song, J.; Wang, G.; and Shen, H. T. 2023 · 2023
Later among the works it cites.
Spatial-information Guided Adaptive Context-aware Network for Efficient RGBD Semantic Segmentation
Zhang, Y.; Xiong, C.; Liu, J.; Ye, X.; and Sun, G. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving Cross-Modal Image-Text Retrieval With Teacher-Student Learning
Liu, J.; Yang, M.; Li, C.; and Xu, R. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
CMPD: Using Cross Memory Network With Pair Discrimination for Image-Text Retrieval
Wen, X.; Han, Z.; and Liu, Y.-S. 2021 · 2021
Cited alongside, same era.
Geometry-aware guided loss for deep crack recognition
Chen, Z.; Zhang, J.; Lai, Z.; Chen, J.; Liu, Z.; and Li, J. 2022 · 2022
Cited alongside, same era.
A Differentiable Semantic Metric Approximation in Probabilistic Embedding for Cross-Modal Retrieval
Li, H.; Song, J.; Gao, L.; Zeng, P.; Zhang, H.; and Li, G. 2022a
Cited in the paper.
Discrete Fusion Adversarial Hashing for cross-modal retrieval
Li, J.; Yu, E.; Ma, J.; Chang, X.; Zhang, H.; and Sun, J. 2022b
Cited in the paper.
Neuron-Based Spiking Transmission and Reasoning Network for Robust Image-Text Retrieval
Li, W.; Ma, Z.; Deng, L.-J.; Fan, X.; and Tian, Y. 2023a
Cited in the paper.
Chu, J.; Li, W.; Wang, X.; Ning, K.; Lu, Y.; and Fan, X. 2024 · 2024
Closest in time.
Multi-layer Probabilistic Association Reasoning Network for Image-Text Retrieval
Li, W.; Xiong, R.; and Fan, X. 2024 · 2024
Closest in time.
Divide and augment: Supervised domain adaptation via sample-wise feature fusion
Chen, Z.; Pu, B.; Zhao, L.; He, J.; and Liang, P. 2025 · 2025
Closest in time.