Fetching the paper…
Reading the bibliography…
We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Stealing machine learning models via prediction APIs
Tramèr, F., Zhang, F., Juels, A., Reiter, M. K., and Ristenpart, T · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Veit, A., Wilber, M. J., and Belongie, S. J · 2016
Earlier work this paper cites.
PRADA: protecting against DNN model stealing attacks
Juuti, M., Szyller, S., Marchal, S., and Asokan, N · 2019
Earlier work this paper cites.
Model reconstruction from model explanations
Milli, S., Schmidt, L., Dragan, A. D., and Hardt, M · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
Cryptanalytic extraction of neural network models
Carlini, N., Jagielski, M., and Mironov, I · 2020
Earlier work this paper cites.
Stateful detection of black-box adversarial attacks
Chen, S., Carlini, N., and Wagner, D · 2020
Earlier work this paper cites.
High accuracy and high fidelity extraction of neural networks
Jagielski, M., Carlini, N., Berthelot, D., Kurakin, A., and Papernot, N · 2020
Earlier work this paper cites.
Reverse-engineering deep relu networks
Rolnick, D. and Kording, K · 2020
Earlier work this paper cites.
Leaky DNN: Stealing deep-learning model secret with GPU context-switching side-channel
Wei, J., Zhang, Y., Zhou, Z., Li, Z., and Al Faruque, M. A · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
On the sizes of OpenAI API models
Gao, L · 2021
Earlier work this paper cites.
Stateful detection of model extraction attacks
Pal, S., Gupta, Y., Kanade, A., and Shevade, S · 2021
Cited alongside, same era.
FUDGE: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Cited alongside, same era.
Grey-box extraction of natural language models
Zanella-Beguelin, S., Tople, S., Paverd, A., and Köpf, B · 2021
Cited alongside, same era.
8-bit Optimizers via Block-wise Quantization
Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L · 2022
Cited alongside, same era.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., and Mordatch, I · 2022
Cited alongside, same era.
Scaling language models: Methods, analysis and insights from training gopher, 2022
Polynomial time cryptanalytic extraction of neural network models
Shamir, A., Canales-Martinez, I., Hambitzer, A., Chavez-Saab, J., Rodrigez-Henriquez, F., and Satpute, N · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Common LLM settings, 2024
Biderman, S · 2024
Closest in time.
Spectral filters, dark signals, and attention sinks, 2024
Cancedda, N · 2024
Closest in time.
openlogprobs, 2024
Chiu, J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2022
Cited alongside, same era.
PaLM 2 Technical Report, 2023
Anil, R. et al · 2023
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Cited alongside, same era.
Privacy side channels in machine learning systems
Debenedetti, E., Severi, G., Carlini, N., Choquette-Choo, C. A., Jagielski, M., Nasr, M., Wallace, E., and Tramèr, F · 2023
Cited alongside, same era.
Stateful defenses for machine learning models are not yet secure against black-box attacks
Feng, R., Hooda, A., Mangaokar, N., Fawaz, K., Jha, S., and Prakash, A · 2023
Cited alongside, same era.
Active retrieval augmented generation
Jiang, Z., Xu, F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G · 2023
Cited alongside, same era.
Morris, J. X., Zhao, W., Chiu, J. T., Shmatikov, V., and Rush, A. M · 2023
Cited alongside, same era.
Finlayson, M., Swayamdipta, S., and Ren, X · 2024
Closest in time.
Changelog 1.38.0, 2024
Google · 2024
Closest in time.
Universal neurons in gpt2 language models, 2024
Gurnee, W., Horsley, T., Guo, Z. C., Kheirkhah, T. R., Sun, Q., Hathaway, W., Nanda, N., and Bertsimas, D · 2024
Closest in time.
Query-based adversarial prompt generation
Hayase, J., Borevkovic, E., Carlini, N., Tramèr, F., and Nasr, M · 2024
Closest in time.
Tuning language models by proxy
Liu, A., Han, X., Wang, Y., Tsvetkov, Y., Choi, Y., and Smith, N. A · 2024
Closest in time.
An emulator for fine-tuning large language models using small language models
Mitchell, E., Rafailov, R., Sharma, A., Finn, C., and Manning, C. D · 2024
Closest in time.
Using logit bias to define token probability, 2023
OpenAI · 2024
Closest in time.
Create chat completion, 2024
OpenAI · 2024
Closest in time.