Fetching the paper…
Reading the bibliography…
Prompt engineering is a technique that involves augmenting a large pre-trained model with task-specific hints, known as prompts, to adapt the model to new tasks.
A computational approach to edge detection
J. Canny · 1986
Earlier work this paper cites.
Dead leaves models: from space tesselation to random functions proc. of the symposium on the advances in the theory and applications of random sets (fontainebleau, 9-11 october 1996) ed d jeulin, 1997
D. Jeulin · 1997
Earlier work this paper cites.
Annealed importance sampling
R. M. Neal · 2001
Earlier work this paper cites.
Learning from experiments when context matters
L. Pritchett and J. Sandefur · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren et al · 2015
Earlier work this paper cites.
Draw: A recurrent neural network for image generation
K. Gregor et al · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein et al · 2015
Earlier work this paper cites.
Generative adversarial text to image synthesis
S. Reed et al · 2016
Earlier work this paper cites.
Variational autoencoder for deep learning of images, labels and captions
Y. Pu et al · 2016
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
P. Isola et al · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry et al · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2018
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford et al · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
A style-based generator architecture for generative adversarial networks
T. Karras et al · 2019
Earlier work this paper cites.
An introduction to variational autoencoders
D. P. Kingma et al · 2019
Earlier work this paper cites.
The global landscape of ai ethics guidelines
A. Jobin et al · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy et al · 2020
Earlier work this paper cites.
The turking test: Can language models understand instructions?
A. Efrat and O. Levy · 2020
Earlier work this paper cites.
AP Art History: 5 Practice Tests + Comprehensive Review + Online Practice
J. B. Nici · 2020
Earlier work this paper cites.
Generative adversarial networks
I. Goodfellow et al · 2020
Earlier work this paper cites.
Nvae: A deep hierarchical variational autoencoder
A. Vahdat and J. Kautz · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho et al · 2020
Earlier work this paper cites.
UNIFIEDQA: Crossing format boundaries with a single QA system
D. Khashabi et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester et al · 2021
Earlier work this paper cites.
X. Liu et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu et al · 2021
Earlier work this paper cites.
Multimodal few-shot learning with frozen language models
M. Tsimpoukelli et al · 2021
Earlier work this paper cites.
Unifying vision-and-language tasks via text generation
J. Cho et al · 2021
Earlier work this paper cites.
Learning how to ask: Querying LMs with mixtures of soft prompts
G. Qin and J. Eisner · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
L. Weidinger et al · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
J. Li et al · 2021
Earlier work this paper cites.
Making Pre-trained Language Models Better Few-shot Learners
T. Gao et al · 2021
Earlier work this paper cites.
Evaluating clip: towards characterization of broader capabilities and downstream implications
S. Agarwal et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh et al · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Earlier work this paper cites.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol et al · 2021
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
B. Mildenhall et al · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Z. Jiang et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester et al · 2021
Earlier work this paper cites.
Few-shot text generation with natural language instructions
T. Schick and H. Schütze · 2021
Earlier work this paper cites.
Template-based named entity recognition using BART
L. Cui et al · 2021
Earlier work this paper cites.
Prompting contrastive explanations for commonsense reasoning tasks
B. Paranjape et al · 2021
Earlier work this paper cites.
What makes good in-context examples for gpt-
J. Liu et al · 2021
Earlier work this paper cites.
Unadversarial examples: Designing objects for robust vision
H. Salman et al · 2021
Earlier work this paper cites.
Effective and efficient vote attack on capsule networks
J. Gu et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac et al · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei et al · 2022
Earlier work this paper cites.
A survey for in-context learning
Q. Dong et al · 2022
Earlier work this paper cites.
Reasoning with language model prompting: A survey
S. Qiao et al · 2022
Earlier work this paper cites.
Exploring visual prompts for adapting large-scale models
H. Bahng et al · 2022
Earlier work this paper cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
Z. Yang et al · 2022
Earlier work this paper cites.
Vision-and-language pretrained models: A survey
S. Long et al · 2022
Earlier work this paper cites.
Flava: A foundational language and vision alignment model
A. Singh et al · 2022
Earlier work this paper cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
P. Wang et al · 2022
Earlier work this paper cites.
SimVLM: Simple visual language model pretraining with weak supervision
Z. Wang et al · 2022
Earlier work this paper cites.
MAGMA – multimodal augmentation of generative models through adapter-based finetuning
C. Eichenberg et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
S. Zhang et al · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
H. W. Chung et al · 2022
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
O. Rubin et al · 2022
Earlier work this paper cites.
Prompt tuning for generative multimodal pretrained models
H. Yang et al · 2022
Earlier work this paper cites.
Unified Vision and Language Prompt Learning
Y. Zang et al · 2022
Earlier work this paper cites.
Threats to pre-trained language models: Survey and taxonomy
S. Guo et al · 2022
Earlier work this paper cites.
Benchmarking robustness under distribution shift of multimodal image-text models
J. Qiu et al · 2022
Earlier work this paper cites.
Unsupervised prompt learning for vision-language models
T. Huang et al · 2022
Earlier work this paper cites.
Test-time prompt tuning for zero-shot generalization in vision-language models
M. Shu et al · 2022
Cited alongside, same era.
Learning to prompt for vision-language models
K. Zhou et al · 2022
Cited alongside, same era.
Prompting visual-language models for efficient video understanding
C. Ju et al · 2022
Cited alongside, same era.
Multitask Vision-Language Prompt Tuning
S. Shen et al · 2022
Cited alongside, same era.
Conditional prompt learning for vision-language models
K. Zhou et al · 2022
Cited alongside, same era.
Visual prompt tuning
M. Jia et al · 2022
Cited alongside, same era.
PaLI: A jointly-scaled multilingual language-image model
X. Chen et al · 2023
Closest in time.
J. Li et al · 2023
Closest in time.
Unified demonstration retriever for in-context learning
X. Li et al · 2023
Closest in time.
Compositional exemplars for in-context learning
J. Ye et al · 2023
Closest in time.
The multimodal and modular ai chef: Complex recipe generation from imagery
D. Noever and S. E. M. Noever · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Wu et al · 2022
Cited alongside, same era.
CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
Y. Yao et al · 2022
Cited alongside, same era.
Visual prompting via image inpainting
A. Bar et al · 2022
Cited alongside, same era.
Dualcoop: Fast adaptation to multi-label recognition with limited annotations
X. Sun et al · 2022
Cited alongside, same era.
Visual Prompt Tuning for Few-Shot Text Classification
J. Wen et al · 2022
Cited alongside, same era.
Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
X. Gu et al · 2022
Cited alongside, same era.
J. Li et al · 2023
Closest in time.
Kosmos-2: Grounding multimodal large language models to the world
Z. Peng et al · 2023
Closest in time.
https://openai.com/blog/chatgpt
Chatgpt · 2023
Closest in time.
K. Zhang et al · 2023
Closest in time.
On evaluating adversarial robustness of large vision-language models
Y. Zhao et al · 2023
Closest in time.
Benchmarking robustness of adaptation methods on pre-trained vision-language models
S. Chen et al · 2023
Closest in time.
Towards robust prompts on vision-language models
J. Gu et al · 2023
Closest in time.
Multi-event video-text retrieval
G. Zhang et al · 2023
Closest in time.
Diversity-aware meta visual prompting
Q. Huang et al · 2023
Closest in time.
What does clip know about a red circle? visual prompt engineering for vlms
A. Shtedritski et al · 2023
Closest in time.
Maple: Multi-modal prompt learning
M. U. Khattak et al · 2023
Closest in time.
LPT: Long-tailed Prompt Tuning for Image Classification
B. Dong et al · 2023
Closest in time.
Texts as images in prompt tuning for multi-label image recognition
Z. Guo et al · 2023
Closest in time.
Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection
K. Gao et al · 2023
Closest in time.
A. Kirillov et al · 2023
Closest in time.
Understanding Zero-Shot Adversarial Robustness for Large-Scale Models, April 2023
C. Mao et al · 2023
Closest in time.
Visual prompting for adversarial robustness
A. Chen et al · 2023
Closest in time.
Effective robustness against natural distribution shifts for models with different training data
Z. Shi et al · 2023
Closest in time.
Cleanclip: Mitigating data poisoning attacks in multimodal contrastive learning
H. Bansal et al · 2023
Closest in time.
Debiasing vision-language models via biased prompts
C.-Y. Chuang et al · 2023
Closest in time.
Mitigating test-time bias for fair image retrieval
F. Kong et al · 2023
Closest in time.
Balancing the picture: Debiasing vision-language datasets with synthetic contrast sets
B. Smith et al · 2023
Closest in time.
Drag your gan: Interactive point-based manipulation on the generative image manifold
X. Pan et al · 2023
Closest in time.
W. Wu et al · 2023
Closest in time.
Imaginarynet: Learning object detectors without real images and annotations
M. Ni et al · 2023
Closest in time.
Is synthetic data from generative models ready for image recognition?
R. He et al · 2023
Closest in time.
An image is worth one word: Personalizing text-to-image generation using textual inversion
R. Gal et al · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
N. Ruiz et al · 2023
Closest in time.
Diffusion self-guidance for controllable image generation
D. Epstein et al · 2023
Closest in time.
Imagic: Text-based real image editing with diffusion models
B. Kawar et al · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
L. Zhang and M. Agrawala · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing
C. Qi et al · 2023
Closest in time.
Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation
L. Ruan et al · 2023
Closest in time.
J. Zhu et al · 2023
Closest in time.
Diffrf: Rendering-guided 3d radiance field diffusion
N. Müller et al · 2023
Closest in time.
Scaling robot learning with semantically imagined experience
T. Yu et al · 2023
Closest in time.
Trace and pace: Controllable pedestrian animation via guided trajectory diffusion
D. Rempe et al · 2023
Closest in time.
Housediffusion: Vector floorplan generation via a diffusion model with discrete and continuous denoising
M. A. Shabani et al · 2023
Closest in time.
Zero-shot generation of coherent storybook from plain text story using diffusion models
H. Jeong et al · 2023
Closest in time.
Multimodal procedural planning via dual text-image prompting
Y. Lu et al · 2023
Closest in time.
Diffusion models for imperceptible and transferable adversarial attack
J. Chen et al · 2023
Closest in time.
A pilot study of query-free adversarial attack against stable diffusion
H. Zhuang et al · 2023
Closest in time.
Text-to-image diffusion models can be easily backdoored through multimodal data poisoning
S. Zhai et al · 2023
Closest in time.
Zero-day backdoor attack against text-to-image diffusion models via personalization
Y. Huang et al · 2023
Closest in time.
Fair diffusion: Instructing text-to-image generation models on fairness
F. Friedrich et al · 2023
Closest in time.
Social biases through the text-to-image generation lens
R. Naik and B. Nushi · 2023
Closest in time.
T2iat: Measuring valence and stereotypical biases in text-to-image generation
J. Wang et al · 2023
Closest in time.
Stable bias: Analyzing societal representations in diffusion models
A. S. Luccioni et al · 2023
Closest in time.
Dear: Debiasing vision-language models with additive residuals
A. Seth et al · 2023
Closest in time.
Explaining visual biases as words by generating captions
Y. Kim et al · 2023
Closest in time.
Are diffusion models vulnerable to membership inference attacks?
J. Duan et al · 2023
Closest in time.
A reproducible extraction of training images from diffusion models
R. Webster · 2023
Closest in time.
Prompt stealing attacks against text-to-image generation models
X. Shen et al · 2023
Closest in time.
R. Anil et al · 2023
Closest in time.
Promptattack: Probing dialogue state trackers with adversarial prompts
X. Dong et al · 2023
Closest in time.
Is prompt all you need? no. A comprehensive and broader view of instruction learning
R. Lou et al · 2023
Closest in time.
Images speak in images: A generalist painter for in-context visual learning
X. Wang et al · 2023
Closest in time.
Prompt generation networks for input-based adaptation of frozen vision transformers
J. Loedeman et al · 2023
Closest in time.
Exploring the benefits of visual prompting in differential privacy
Y. Li et al · 2023
Closest in time.
Imagebind: One embedding space to bind them all
R. Girdhar et al · 2023
Closest in time.
Otter: A multi-modal model with in-context instruction tuning
B. Li et al · 2023
Closest in time.
Llama-adapter v2: Parameter-efficient visual instruction model
P. Gao et al · 2023
Closest in time.
Backdoor defense via adaptively splitting poisoned dataset
K. Gao et al · 2023
Closest in time.
Backdoor defense via suppressing model shortcuts
S. Yang et al · 2023
Closest in time.
Frontier ai regulation: Managing emerging risks to public safety
M. Anderljung et al · 2023
Closest in time.
Regulating chatgpt and other large generative ai models
P. Hacker et al · 2023
Closest in time.
Do dall-e and flamingo understand each other?
H. Li et al · 2023
Closest in time.