Fetching the paper…
Reading the bibliography…
Research connecting text and images has recently seen several breakthroughs, with models like CLIP, DALL-E 2, and Stable Diffusion.
Content based image retrieval: survey
Mehwish Rehman, Muhammad Iqbal, Muhammad Sharif, and Mudassar Raza · 2012
Earlier work this paper cites.
Personalised information retrieval: survey and classification
M Rami Ghorab, Dong Zhou, Alexander O’connor, and Vincent Wade · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Earlier work this paper cites.
Weather influence and classification with automotive lidar sensors
Robin Heinzler, Philipp Schindler, Jürgen Seekircher, Werner Ritter, and Wilhelm Stork · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
Dave Zhenyu Chen, Angel X Chang, and Matthias Nießner · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
A survey on 3d lidar localization for autonomous vehicles
Mahdi Elhousni and Xinming Huang · 2020
Earlier work this paper cites.
Recent advances in open set recognition: A survey
Chuanxing Geng, Sheng-jun Huang, and Songcan Chen · 2020
Earlier work this paper cites.
Weather classification using an automotive lidar sensor based on detections on asphalt and atmosphere
Jose Roberto Vargas Rivero, Thiemo Gerbich, Valentina Teiluf, Boris Buschardt, and Jia Chen · 2020
Earlier work this paper cites.
4d panoptic lidar segmentation
Mehmet Aygun, Aljosa Osep, Mark Weber, Maxim Maximov, Cyrill Stachniss, Jens Behley, and Laura Leal-Taixé · 2021
Earlier work this paper cites.
Diffusion models beat GANs on image synthesis
Prafulla Dhariwal and Alexander Quinn Nichol · 2021
Earlier work this paper cites.
One million scenes for autonomous driving: ONCE dataset
Jiageng Mao, Minzhe Niu, Chenhan Jiang, hanxue liang, Jingheng Chen, Xiaodan Liang, Yamin Li, Chaoqiang Ye, Wei Zhang, Zhenguo Li, Jie Yu, Hang Xu, and Chunjing Xu · 2021
Earlier work this paper cites.
Generative zero-shot learning for semantic segmentation of 3d point clouds
Björn Michele, Alexandre Boulch, Gilles Puy, Maxime Bucher, and Renaud Marlet · 2021
Earlier work this paper cites.
Joint passage ranking for diverse multi-answer retrieval
Sewon Min, Kenton Lee, Ming-Wei Chang, Kristina Toutanova, and Hannaneh Hajishirzi · 2021
Earlier work this paper cites.
Clipcap: Clip prefix for image captioning
Ron Mokady, Amir Hertz, and Amit H Bermano · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Part2word: Learning joint embedding of point clouds and text by matching parts to words
Chuan Tang, Xi Yang, Bojian Wu, Zhizhong Han, and Yi Chang · 2021
Cited alongside, same era.
Deep 3d object detection networks using lidar data: A review
Yutian Wu, Yueyu Wang, Shuwei Zhang, and Harutoshi Ogai · 2021
Cited alongside, same era.
VideoCLIP: Contrastive pre-training for
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer · 2021
Cited alongside, same era.
Semantic segmentation of 3d lidar data using deep learning: a review of projection-based methods
Alok Jhaldiyal and Navendu Chaudhary · 2022
Closest in time.
Image segmentation using text and image prompts
Timo Lüddecke and Alexander Ecker · 2022
Closest in time.
Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li · 2022
Closest in time.
Ei-clip: Entity-aware interventional contrastive learning for e-commerce cross-modal retrieval
Haoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale, Zhangyang Wang, Tong Yu, Jiuxiang Gu, Sunav Choudhary, and Xiaohui Xie · 2022
Closest in time.
X-clip: End-to-end multi-grained contrastive learning for video-text retrieval
Yiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan, Ji Zhang, and Rongrong Ji · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
3dvg-transformer: Relation modeling for visual grounding on point clouds
Lichen Zhao, Daigang Cai, Lu Sheng, and Dong Xu · 2021
Cited alongside, same era.
Effective conditioned and composed image retrieval combining clip-based features
Alberto Baldrati, Marco Bertini, Tiberio Uricchio, and Alberto Del Bimbo · 2022
Cited alongside, same era.
Cross-lingual and multilingual clip
Fredrik Carlsson, Philipp Eisen, Faton Rekathati, and Magnus Sahlgren · 2022
Cited alongside, same era.
Zero-shot learning on 3d point cloud objects and beyond
Ali Cheraghian, Shafin Rahman, Townim F Chowdhury, Dylan Campbell, and Lars Petersson · 2022
Cited alongside, same era.
Embracing single stride 3d object detector with sparse transformer
Lue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang, Hang Zhao, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang · 2022
Cited alongside, same era.
Embracing single stride 3d object detector with sparse transformer
Lue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang, Hang Zhao, Feng Wang, Naiyan Wang, and Zhaoxiang Zhang · 2022
Cited alongside, same era.
Ensemble deep learning: A review
M.A. Ganaie, Minghui Hu, A.K. Malik, M. Tanveer, and P.N. Suganthan · 2022
Cited alongside, same era.
Gaurav Parmar et al · 2022
Closest in time.
Hierarchical text-conditional image generation with CLIP latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Closest in time.
Language-grounded indoor 3d semantic segmentation in the wild
David Rozenberszki, Or Litany, and Angela Dai · 2022
Closest in time.
Clip-forge: Towards zero-shot text-to-shape generation
Aditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, and Kamal Rahimi Malekshan · 2022
Closest in time.
Clip-nerf: Text-and-image driven manipulation of neural radiance fields
Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao · 2022
Closest in time.
Cris: Clip-driven referring image segmentation
Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu · 2022
Closest in time.
Wav2clip: Learning robust audio representations from clip
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello · 2022
Closest in time.
Pointclip: Point cloud understanding by CLIP
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li · 2022
Closest in time.
Extract free dense labels from clip
Chong Zhou, Chen Change Loy, and Bo Dai · 2022
Closest in time.
The role of imagenet classes in fréchet inception distance
Tuomas Kynkäänniemi, Tero Karras, Miika Aittala, Timo Aila, and Jaakko Lehtinen · 2023
Closest in time.