Fetching the paper…
Reading the bibliography…
The era of vision-language models (VLMs) trained on web-scale datasets challenges conventional formulations of "open-world" perception.
“Towards long-tailed 3d detection”
Neehar Peri, Achal Dave, Deva Ramanan and Shu Kong · 1915
Earlier work this paper cites.
“ImageNet: A large-scale hierarchical image database”
Jia Deng et al · 2009
Earlier work this paper cites.
“The Pascal Visual Object Classes (VOC) Challenge”
M. Everingham et al · 2010
Earlier work this paper cites.
“Microsoft COCO: Common Objects in Context”
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
“Overcoming catastrophic forgetting in neural networks”
James Kirkpatrick et al · 2017
Earlier work this paper cites.
“Growing a brain: Fine-tuning by increasing model capacity”
Yu-Xiong Wang, Deva Ramanan and Martial Hebert · 2017
Earlier work this paper cites.
“Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning”
Piyush Sharma, Nan Ding, Sebastian Goodman and Radu Soricut · 2018
Earlier work this paper cites.
“LVIS: A dataset for large vocabulary instance segmentation”
Agrim Gupta, Piotr Dollar and Ross Girshick · 2019
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP”
Neil Houlsby et al · 2019
Earlier work this paper cites.
“Few-shot object detection via feature reweighting”
Bingyi Kang et al · 2019
Earlier work this paper cites.
“Objects365: A large-scale, high-quality dataset for object detection”
Shuai Shao et al · 2019
Earlier work this paper cites.
“Meta r-cnn: Towards general solver for instance-level low-shot learning”
Xiaopeng Yan et al · 2019
Earlier work this paper cites.
“nuscenes: A multimodal dataset for autonomous driving”
Holger Caesar et al · 2020
Earlier work this paper cites.
“A simple framework for contrastive learning of visual representations”
Ting Chen, Simon Kornblith, Mohammad Norouzi and Geoffrey Hinton · 2020
Earlier work this paper cites.
“Few-shot object detection with attention-RPN and multi-relation detector”
Qi Fan, Wei Zhuo, Chi-Keung Tang and Yu-Wing Tai · 2020
Earlier work this paper cites.
“Making pre-trained language models better few-shot learners”
Tianyu Gao, Adam Fisch and Danqi Chen · 2020
Earlier work this paper cites.
“Momentum contrast for unsupervised visual representation learning”
Kaiming He et al · 2020
Earlier work this paper cites.
“How can we know what language models know?”
Zhengbao Jiang, Frank Xu, Jun Araki and Graham Neubig · 2020
Earlier work this paper cites.
“Autoprompt: Eliciting knowledge from language models with automatically generated prompts”
Taylor Shin et al · 2020
Earlier work this paper cites.
“Frustratingly Simple Few-Shot Object Detection”
Xin Wang et al · 2020
Earlier work this paper cites.
“Multi-scale positive sample refinement for few-shot object detection”
Jiaxi Wu, Songtao Liu, Di Huang and Yunhong Wang · 2020
Earlier work this paper cites.
“Side-tuning: a baseline for network adaptation via additive side networks”
Jeffrey Zhang et al · 2020
Earlier work this paper cites.
“Generalized few-shot object detection without forgetting”
Zhibo Fan, Yuchen Ma, Zeming Li and Jian Sun · 2021
Earlier work this paper cites.
“Open-vocabulary object detection via vision and language knowledge distillation”
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo and Yin Cui · 2021
Cited alongside, same era.
“Open-vocabulary object detection via vision and language knowledge distillation”
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo and Yin Cui · 2021
Cited alongside, same era.
“BERTese: Learning to speak to BERT”
Adi Haviv, Jonathan Berant and Amir Globerson · 2021
Cited alongside, same era.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Cited alongside, same era.
“A Spoken Language Dataset of Descriptions for Speech-Based Grounded Language Learning”
Gaoussou Kebe et al · 2021
“MedCLIP: Contrastive Learning from Unpaired Medical Images and Text”
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal and Jimeng Sun · 2022
Later among the works it cites.
“Few-shot object detection and viewpoint estimation for objects in the wild”
Yang Xiao, Vincent Lepetit and Renaud Marlet · 2022
Later among the works it cites.
“Regionclip: Region-based language-image pretraining”
Yiwu Zhong et al · 2022
Later among the works it cites.
“Conditional prompt learning for vision-language models”
Kaiyang Zhou, Jingkang Yang, Chen Loy and Ziwei Liu · 2022
Later among the works it cites.
“Learning to prompt for vision-language models”
Kaiyang Zhou, Jingkang Yang, Chen Loy and Ziwei Liu · 2022
Later among the works it cites.
“Detecting twenty-thousand classes using image-level supervision”
Xingyi Zhou et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Few-shot object detection: A comprehensive survey”
Mona Köhler, Markus Eisenbach and Horst-Michael Gross · 2021
Cited alongside, same era.
“The power of scale for parameter-efficient prompt tuning”
Brian Lester, Rami Al-Rfou and Noah Constant · 2021
Cited alongside, same era.
“Beyond max-margin: Class margin equilibrium for few-shot object detection”
Bohao Li et al · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Cited alongside, same era.
“Fsce: Few-shot object detection via contrastive proposal encoding”
Bo Sun et al · 2021
Cited alongside, same era.
“Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting”
Benjamin Wilson et al · 2021
Cited alongside, same era.
“Universal-prototype enhancing for few-shot object detection”
Aming Wu, Yahong Han, Linchao Zhu and Yi Yang · 2021
Cited alongside, same era.
Josh Achiam et al · 2023
Closest in time.
“Thinking Like an Annotator: Generation of Dataset Labeling Instructions”
Nadine Chang et al · 2023
Closest in time.
“Grounding dino: Marrying dino with grounded pre-training for open-set object detection”
Shilong Liu et al · 2023
Closest in time.
“DiGeo: Discriminative Geometry-Aware Learning for Generalized Few-Shot Object Detection”
Jiawei Ma et al · 2023
Closest in time.
“Long-Tailed 3D Detection via 2D Late Fusion”, 2023
Yechi Ma et al · 2023
Closest in time.
“Visual Classification via Description from Large Language Models”
Sachit Menon and Carl Vondrick · 2023
Closest in time.
“Scaling Open-Vocabulary Object Detection”
Matthias Minderer, Alexey Gritsenko and Neil Houlsby · 2023
Closest in time.
“Prompting scientific names for zero-shot species recognition”
Shubham Parashar, Zhiqiu Lin, Yanan Li and Shu Kong · 2023
Closest in time.
“Revisiting classifier: Transferring vision-language models for video recognition”
Wenhao Wu, Zhun Sun and Wanli Ouyang · 2023
Closest in time.
“Dual modality prompt tuning for vision-language pre-trained model”
Yinghui Xing et al · 2023
Closest in time.
“Generating Features with Increased Crop-related Diversity for Few-Shot Object Detection”
Jingyi Xu, Hieu Le and Dimitris Samaras · 2023
Closest in time.
“Adding conditional control to text-to-image diffusion models”
Lvmin Zhang, Anyi Rao and Maneesh Agrawala · 2023
Closest in time.
“Prompt-aligned gradient for prompt tuning”
Beier Zhu et al · 2023
Closest in time.
“Clip-adapter: Better vision-language models with feature adapters”
Peng Gao et al · 2024
Closest in time.
“Few-Shot Recognition via Stage-Wise Augmented Finetuning”
Tian Liu, Huixin Zhang, Shubham Parashar and Shu Kong · 2024
Closest in time.
“The Neglected Tails in Vision-Language Models”
Shubham Parashar et al · 2024
Closest in time.
“Multi-modal queried object detection in the wild”
Yifan Xu et al · 2024
Closest in time.