Fetching the paper…
Reading the bibliography…
Visual prompt engineering is a fundamental technology in the field of visual and image Artificial General Intelligence, serving as a key component for achieving zero-shot capabilities.
Mind children: The future of robot and human intelligence
Hans Moravec · 1988
Earlier work this paper cites.
Elephants don’t play chess
Rodney A Brooks · 1990
Earlier work this paper cites.
Align: a program to superimpose protein coordinates, accounting for insertions and deletions
Gerson H Cohen · 1997
Earlier work this paper cites.
Web search intent induction via automatic query reformulation
Hal Daumé III and Eric Brill · 2004
Earlier work this paper cites.
Image segmentation with a bounding box prior
Victor Lempitsky, Pushmeet Kohli, Carsten Rother, and Toby Sharp · 2009
Earlier work this paper cites.
icoseg: Interactive co-segmentation with intelligent scribble guidance
Dhruv Batra, Adarsh Kowdle, Devi Parikh, Jiebo Luo, and Tsuhan Chen · 2010
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Earlier work this paper cites.
Milcut: A sweeping line multiple instance learning paradigm for interactive image segmentation
Jiajun Wu, Yibiao Zhao, Jun-Yan Zhu, Siwei Luo, and Zhuowen Tu · 2014
Earlier work this paper cites.
Error-tolerant scribbles based interactive image segmentation
Junjie Bai and Xiaodong Wu · 2014
Earlier work this paper cites.
Deep interactive object selection
Ning Xu, Brian Price, Scott Cohen, Jimei Yang, and Thomas S Huang · 2016
Earlier work this paper cites.
Deepcut: Object segmentation from bounding box annotations using convolutional neural networks
Martin Rajchl, Matthew CH Lee, Ozan Oktay, Konstantinos Kamnitsas, Jonathan Passerat-Palmbach, Wenjia Bai, Mellisa Damodaram, Mary A Rutherford, Joseph V Hajnal, Bernhard Kainz, et al · 2016
Earlier work this paper cites.
Scribblesup: Scribble-supervised convolutional networks for semantic segmentation
Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Ask the right questions: Active question reformulation with reinforcement learning
Christian Buck, Jannis Bulian, Massimiliano Ciaramita, Wojciech Gajewski, Andrea Gesmundo, Neil Houlsby, and Wei Wang · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Earlier work this paper cites.
One-shot learning for semantic segmentation
Amirreza Shaban, Shray Bansal, Zhen Liu, Irfan Essa, and Byron Boots · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Interactive medical image segmentation using deep learning with image-specific fine tuning
Guotai Wang, Wenqi Li, Maria A Zuluaga, Rosalind Pratt, Premal A Patel, Michael Aertsen, Tom Doel, Anna L David, Jan Deprest, Sébastien Ourselin, et al · 2018
Earlier work this paper cites.
Few-shot semantic segmentation with prototype learning
Nanqing Dong and Eric P Xing · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Commongen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Interactive image segmentation via backpropagating refinement scheme
Won-Dong Jang and Chang-Su Kim · 2019
Earlier work this paper cites.
Panet: Few-shot image semantic segmentation with prototype alignment
Kaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou, and Jiashi Feng · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Xcopa: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen · 2020
Earlier work this paper cites.
Few-shot text generation with pattern-exploiting training
Timo Schick and Hinrich Schütze · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Interactive image segmentation with first click attention
Zheng Lin, Zhao Zhang, Lin-Zhuo Chen, Ming-Ming Cheng, and Shao-Ping Lu · 2020
Earlier work this paper cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Earlier work this paper cites.
Psa-net: Deep learning–based physician style–aware segmentation network for postoperative prostate cancer clinical target volumes
Anjali Balagopal, Howard Morgan, Michael Dohopolski, Ramsey Timmerman, Jie Shan, Daniel F Heitjan, Wei Liu, Dan Nguyen, Raquibul Hannan, Aurelie Garant, et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Earlier work this paper cites.
Transformer in transformer
Kai Han, An Xiao, Enhua Wu, Jianyuan Guo, Chunjing Xu, and Yunhe Wang · 2021
Earlier work this paper cites.
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He · 2021
Earlier work this paper cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Earlier work this paper cites.
Open-vocabulary object detection via vision and language knowledge distillation
Xiuye Gu, Tsung-Yi Lin, Weicheng Kuo, and Yin Cui · 2021
Earlier work this paper cites.
Actionclip: A new paradigm for video action recognition
Mengmeng Wang, Jiazheng Xing, and Yong Liu · 2021
Earlier work this paper cites.
Conditional diffusion for interactive segmentation
Xi Chen, Zhiyan Zhao, Feiwu Yu, Yilei Zhang, and Manni Duan · 2021
Earlier work this paper cites.
Clinicalradiobert: Knowledge-infused few shot learning for clinical notes named entity recognition
Saed Rezayi, Haixing Dai, Zhengliang Liu, Zihao Wu, Akarsh Hebbar, Andrew H Burns, Lin Zhao, Dajiang Zhu, Quanzheng Li, Wei Liu, et al · 2022
Earlier work this paper cites.
Agribert: knowledge-infused agricultural language models for matching food and nutrition
Saed Rezayi, Zhengliang Liu, Zihao Wu, Chandra Dhakal, Bao Ge, Chen Zhen, Tianming Liu, and Sheng Li · 2022
Earlier work this paper cites.
Swin transformer v2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al · 2022
Earlier work this paper cites.
Mask-guided vision transformer (mg-vit) for few-shot learning
Yuzhong Chen, Zhenxiang Xiao, Lin Zhao, Lu Zhang, Haixing Dai, David Weizhong Liu, Zihao Wu, Changhe Li, Tuo Zhang, Changying Li, et al · 2022
Earlier work this paper cites.
A unified and biologically-plausible relational graph representation of vision transformers
Yuzhong Chen, Yu Du, Zhenxiang Xiao, Lin Zhao, Lu Zhang, David Weizhong Liu, Dajiang Zhu, Tuo Zhang, Xintao Hu, Tianming Liu, et al · 2022
Earlier work this paper cites.
Rectify vit shortcut learning by visual saliency
Chong Ma, Lin Zhao, Yuzhong Chen, David Weizhong Liu, Xi Jiang, Tuo Zhang, Xintao Hu, Dinggang Shen, Dajiang Zhu, and Tianming Liu · 2022
Earlier work this paper cites.
Classification of alzheimer’s disease via vision transformer: Classification of alzheimer’s disease via vision transformer
Yanjun Lyu, Xiaowei Yu, Dajiang Zhu, and Lu Zhang · 2022
Earlier work this paper cites.
Disentangling spatial-temporal functional brain networks via twin-transformers
Xiaowei Yu, Lu Zhang, Lin Zhao, Yanjun Lyu, Tianming Liu, and Dajiang Zhu · 2022
Earlier work this paper cites.
Accurate and efficient deep neural network based deformable image registration method in lung cancer
Y Ding, Z Liu, H Feng, J Holmes, Y Yang, N Yu, T Sio, S Schild, B Li, and W Liu · 2022
Cited alongside, same era.
Discovering dynamic functional brain networks via spatial and channel-wise attention
Yiheng Liu, Enjie Ge, Mengshen He, Zhengliang Liu, Shijie Zhao, Xintao Hu, Dajiang Zhu, Tianming Liu, and Bao Ge · 2022
Cited alongside, same era.
Reducing retraining by recycling parameter-efficient prompts
Brian Lester, Joshua Yurtsever, Siamak Shakeri, and Noah Constant · 2022
Cited alongside, same era.
Coarse-to-fine knowledge graph domain adaptation based on distantly-supervised iterative training
Homgmin Cai, Wenxiong Liao, Zhengliang Liu, Xiaoke Huang, Yiyang Zhang, Siqi Ding, Sheng Li, Quanzheng Li, Tianming Liu, and Xiang Li · 2022
Cited alongside, same era.
All in one: Exploring unified video-language pre-training
Jinpeng Wang, Yixiao Ge, Rui Yan, Yuying Ge, Kevin Qinghong Lin, Satoshi Tsutsui, Xudong Lin, Guanyu Cai, Jianping Wu, Ying Shan, et al · 2023
Closest in time.
Community graph convolution neural network for alzheimer’s disease classification and pathogenetic factors identification
Xia-An Bi, Ke Chen, Siyu Jiang, Sheng Luo, Wenyan Zhou, Zhaoxu Xing, Luyun Xu, Zhengliang Liu, and Tianming Liu · 2023
Closest in time.
Lian Zhang, Jason M Holmes, Zhengliang Liu, Sujay A Vora, Terence T Sio, Carlos E Vargas, Nathan Y Yu, Sameer R Keole, Steven E Schild, Martin Bues, et al · 2023
Closest in time.
Deep-learning based fast and accurate 3d ct deformable image registration in lung cancer
Yuzhen Ding, Hongying Feng, Yunze Yang, Jason Holmes, Zhengliang Liu, David Liu, William W Wong, Nathan Y Yu, Terence T Sio, Steven E Schild, et al · 2023
Closest in time.
Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei · 2022
Cited alongside, same era.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al · 2022
Cited alongside, same era.
Transformers in vision: A survey
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Cited alongside, same era.
Dezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin, and Yang Liu · 2023
Closest in time.
Artificial general intelligence for medical imaging
Xiang Li, Lu Zhang, Zihao Wu, Zhengliang Liu, Lin Zhao, Yixuan Yuan, Jun Liu, Gang Li, Dajiang Zhu, Pingkuan Yan, et al · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
How segment anything model (sam) boost medical image segmentation?
Yichi Zhang and Rushi Jiao · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Closest in time.
Image as a foreign language: Beit pretraining for vision and vision-language tasks
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al · 2023
Closest in time.
Diffusion models in vision: A survey
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah · 2023
Closest in time.
Open-vocabulary semantic segmentation with mask-adapted clip
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu · 2023
Closest in time.
Visual prompt tuning for generative transfer learning
Kihyuk Sohn, Huiwen Chang, José Lezama, Luisa Polania, Han Zhang, Yuan Hao, Irfan Essa, and Lu Jiang · 2023
Closest in time.
Ruining Deng, Can Cui, Quan Liu, Tianyuan Yao, Lucas W Remedios, Shunxing Bao, Bennett A Landman, Lee E Wheless, Lori A Coburn, Keith T Wilson, et al · 2023
Closest in time.
Segment anything model for medical image analysis: an experimental study
Maciej A Mazurowski, Haoyu Dong, Hanxue Gu, Jichen Yang, Nicholas Konz, and Yixin Zhang · 2023
Closest in time.
Medical sam adapter: Adapting segment anything model for medical image segmentation
Junde Wu, Rao Fu, Huihui Fang, Yuanpei Liu, Zhaowei Wang, Yanwu Xu, Yueming Jin, and Tal Arbel · 2023
Closest in time.
Tao Zhou, Yizhe Zhang, Yi Zhou, Ye Wu, and Chen Gong · 2023
Closest in time.
Generalist vision foundation models for medical imaging: A case study of segment anything model on zero-shot medical segmentation
Peilun Shi, Jianing Qiu, Sai Mu Dalike Abaxi, Hao Wei, Frank P-W Lo, and Wu Yuan · 2023
Closest in time.
Accuracy of segment-anything model (sam) in medical image segmentation tasks
Sheng He, Rina Bao, Jingpeng Li, P Ellen Grant, and Yangming Ou · 2023
Closest in time.
Input augmentation with sam: Boosting medical image segmentation with segmentation foundation model
Yizhe Zhang, Tao Zhou, Peixian Liang, and Danny Z Chen · 2023
Closest in time.
Yangming Cheng, Liulei Li, Yuanyou Xu, Xiaodi Li, Zongxin Yang, Wenguan Wang, and Yi Yang · 2023
Closest in time.
Track anything: Segment anything meets videos
Jinyu Yang, Mingqi Gao, Zhe Li, Shang Gao, Fangjing Wang, and Feng Zheng · 2023
Closest in time.
Automated movement tracking of young autistic children during free play is correlated with clinical features associated with autism
Andrew Yuan, Maura Sabatos-DeVito, Alexandra L Bey, Samantha Major, Kimberly LH Carpenter, Lauren Franz, Jill Howard, Saritha Vermeer, Ryan Simmons, Jesse Troy, et al · 2023
Closest in time.
Ntire 2023 challenge on 360deg omnidirectional image and video super-resolution: Datasets, methods and results
Mingdeng Cao, Chong Mou, Fanghua Yu, Xintao Wang, Yinqiang Zheng, Jian Zhang, Chao Dong, Gen Li, Ying Shan, Radu Timofte, et al · 2023
Closest in time.
Knowledge distillation with segment anything (sam) model for planetary geological mapping
Sahib Julka and Michael Granitzer · 2023
Closest in time.
Scalable mask annotation for video text spotting
Haibin He, Jing Zhang, Mengyang Xu, Juhua Liu, Bo Du, and Dacheng Tao · 2023
Closest in time.
Chunming He, Kai Li, Yachao Zhang, Guoxia Xu, Longxiang Tang, Yulun Zhang, Zhenhua Guo, and Xiu Li · 2023
Closest in time.
Anything-3d: Towards single-view anything reconstruction in the wild
Qiuhong Shen, Xingyi Yang, and Xinchao Wang · 2023
Closest in time.
Sam meets robotic surgery: An empirical study in robustness perspective
An Wang, Mobarakol Islam, Mengya Xu, Yang Zhang, and Hongliang Ren · 2023
Closest in time.
Robot based transurethral bladder tumor resection with automatic detection of tumor cells
Vicente García Díaz, R Dinesh Jackson Samuel, Adhiyaman Manickam, Vijayalakshmi Saravanan, Ashish Kr Luhach, and Sujatha Krishnamoorthy · 2023
Closest in time.
Analyzing schedule dependency and sequencing changes for robotic construction using graph analysis
Tessa Beauchat, Yuqing Hu, Robert M Leicht, and Clinton Suanico · 2023
Closest in time.
Inpaint anything: Segment anything meets image inpainting
Tao Yu, Runseng Feng, Ruoyu Feng, Jinming Liu, Xin Jin, Wenjun Zeng, and Zhibo Chen · 2023
Closest in time.
Sam. md: Zero-shot medical image segmentation capabilities of the segment anything model
Saikat Roy, Tassilo Wald, Gregor Koehler, Maximilian R Rokuss, Nico Disch, Julius Holzschuh, David Zimmerer, and Klaus H Maier-Hein · 2023
Closest in time.
Maple: Multi-modal prompt learning
Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, and Fahad Shahbaz Khan · 2023
Closest in time.
Imagic: Text-based real image editing with diffusion models
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani · 2023
Closest in time.
Galip: Generative adversarial clips for text-to-image synthesis
Ming Tao, Bing-Kun Bao, Hao Tang, and Changsheng Xu · 2023
Closest in time.
Position-guided text prompt for vision-language pre-training
Jinpeng Wang, Pan Zhou, Mike Zheng Shou, and Shuicheng Yan · 2023
Closest in time.
Visual prompt multi-modal tracking
Jiawen Zhu, Simiao Lai, Xin Chen, Dong Wang, and Huchuan Lu · 2023
Closest in time.
Diversity-aware meta visual prompting
Qidong Huang, Xiaoyi Dong, Dongdong Chen, Weiming Zhang, Feifei Wang, Gang Hua, and Nenghai Yu · 2023
Closest in time.
Oneformer: One transformer to rule universal image segmentation
Jitesh Jain, Jiachen Li, Mang Tik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi · 2023
Closest in time.
Seggpt: Segmenting everything in context
Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, and Tiejun Huang · 2023
Closest in time.
Segment everything everywhere all at once
Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee · 2023
Closest in time.
Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks
Hao Li, Jinguo Zhu, Xiaohu Jiang, Xizhou Zhu, Hongsheng Li, Chun Yuan, Xiaohua Wang, Yu Qiao, Xiaogang Wang, Wenhai Wang, et al · 2023
Closest in time.
Segment anything is not always perfect: An investigation of sam on different real-world applications
Wei Ji, Jingjing Li, Qi Bi, Wenbo Li, and Li Cheng · 2023
Closest in time.
Can sam count anything? an empirical study on sam counting
Zhiheng Ma, Xiaopeng Hong, and Qinnan Shangguan · 2023
Closest in time.
Text2seg: Remote sensing image semantic segmentation via text-guided visual foundation models
Jielu Zhang, Zhongliang Zhou, Gengchen Mai, Lan Mu, Mengxuan Hu, and Sheng Li · 2023
Closest in time.
Scaling-up remote sensing segmentation dataset with segment anything model
Di Wang, Jing Zhang, Bo Du, Dacheng Tao, and Liangpei Zhang · 2023
Closest in time.
Tianrun Chen, Lanyun Zhu, Chaotao Ding, Runlong Cao, Shangzhan Zhang, Yan Wang, Zejian Li, Lingyun Sun, Papa Mao, and Ying Zang · 2023
Closest in time.
Caption anything: Interactive image description with diverse multimodal controls
Teng Wang, Jinrui Zhang, Junjie Fei, Yixiao Ge, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao, Ying Shan, et al · 2023
Closest in time.
Segment any anomaly without training via hybrid prompt regularization
Yunkang Cao, Xiaohao Xu, Chen Sun, Yuqi Cheng, Zongwei Du, Liang Gao, and Weiming Shen · 2023
Closest in time.
Edit everything: A text-guided generative system for images editing
Defeng Xie, Ruichen Wang, Jian Ma, Chen Chen, Haonan Lu, Dong Yang, Fobo Shi, and Xiaodong Lin · 2023
Closest in time.
Explain any concept: Segment anything meets concept-based explanation
Ao Sun, Pingchuan Ma, Yuanyuan Yuan, and Shuai Wang · 2023
Closest in time.
Towards agi in computer vision: Lessons learned from gpt and large language models
Lingxi Xie, Longhui Wei, Xiaopeng Zhang, Kaifeng Bi, Xiaotao Gu, Jianlong Chang, and Qi Tian · 2023
Closest in time.
Segment anything in medical images
Jun Ma and Bo Wang · 2023
Closest in time.
Guoyu Lu, Sheng Li, Gengchen Mai, Jin Sun, Dajiang Zhu, Lilong Chai, Haijian Sun, Xianqiao Wang, Haixing Dai, Ninghao Liu, et al · 2023
Closest in time.
Xiao Yang, Haixing Dai, Zihao Wu, Ramesh Bist, Sachin Subedi, Jin Sun, Guoyu Lu, Changying Li, Tianming Liu, and Lilong Chai · 2023
Closest in time.