Fetching the paper…
Reading the bibliography…
Model Leeching is a novel extraction attack targeting Large Language Models (LLMs), capable of distilling task-specific knowledge from a target LLM into a reduced parameter model.
Squad: 100,000+ questions for machine comprehension of text, 2016
2016
Earlier work this paper cites.
Stealing machine learning models via prediction APIs
2016
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems, 2017
2017
Earlier work this paper cites.
Membership inference attacks against machine learning models, 2017
2017
Earlier work this paper cites.
Attention is all you need, 2017
2017
Earlier work this paper cites.
Adversarial attacks and defences: A survey, 2018
2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
2019
Earlier work this paper cites.
The security of machine learning in an adversarial setting: A survey
2019
Earlier work this paper cites.
The turking test: Can language models understand instructions?, 2020
2020
Earlier work this paper cites.
Deepsniffer: A dnn model extraction framework based on learning architectural hints
2020
Cited alongside, same era.
Thieves on sesame street! model extraction of bert-based apis, 2020
2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations, 2020
2020
Cited alongside, same era.
Extracting training data from large language models, 2021
2021
Cited alongside, same era.
Dawn: Dynamic adversarial watermarking of neural networks, 2021
2021
Cited alongside, same era.
Reframing instructional prompts to GPTk’s language
2022
Cited alongside, same era.
Ai as agency without intelligence: on chatgpt, large language models, and other generative models
2023
Closest in time.
Pinch: An adversarial extraction attack framework for deep learning models, 2023
2023
Closest in time.
MITRE ATLAS Adversarial Attack Knowledge Base, 2023
2023
Closest in time.
I know what you trained last summer: A survey on stealing machine learning models and defences
2023
Closest in time.
gpt4all.io, 2023
2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
2023
Closest in time.
A prompt pattern catalog to enhance prompt engineering with chatgpt, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-instruct: Aligning language model with self generated instructions, 2022
2022
Cited alongside, same era.
About Bard
2023
Cited alongside, same era.
Sagemaker data labeling pricing
2023
Cited alongside, same era.
Machine generated text: A comprehensive survey of threat models and detection methods, 2023
2023
Cited alongside, same era.
The limitations of deep learning in adversarial settings
Cited in the paper.
2023
Closest in time.
A survey of large language models, 2023
2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models, 2023
2023
Closest in time.