Fetching the paper…
Reading the bibliography…
The community explored to build private inference frameworks for transformer-based large language models (LLMs) in a server-client setting, where the server holds the model parameters and the client inputs its private data (or prompt) for inference.
Cognitron: A self-organizing multilayered neural network
Kunihiko Fukushima · 1975
Earlier work this paper cites.
How to share a secret
Adi Shamir · 1979
Earlier work this paper cites.
Efficient oblivious transfer protocols
Moni Naor and Benny Pinkas · 2001
Earlier work this paper cites.
Extending oblivious transfers efficiently
Yuval Ishai, Joe Kilian, Kobbi Nissim, and Erez Petrank · 2003
Earlier work this paper cites.
Somewhat practical fully homomorphic encryption
Junfeng Fan and Frederik Vercauteren · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Actively secure ot extension with optimal overhead
Marcel Keller, Emmanuela Orsini, and Peter Scholl · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Cryptonets: Applying neural networks to encrypted data with high throughput and accuracy
Nathan Dowlin, Ran Gilad-Bachrach, Kim Laine, Kristin Lauter, Michael Naehrig, and John Wernsing · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
A full rns variant of fv like somewhat homomorphic encryption schemes
Jean Claude Bajard, Julien Eynard, M. Hasan, and Vincent Zucca · 2017
Earlier work this paper cites.
Oblivious neural network predictions via minionn transformations
Jian Liu, Mika Juuti, Yao Lu, and N. Asokan · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
GAZELLE: A low latency framework for secure neural network inference
Chiraag Juvekar, Vinod Vaikuntanathan, and Anantha Chandrakasan · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
Crypten: Secure multi-party computation meets machine learning
Brian Knott, Shobha Venkataraman, Awni Hannun, Shubho Sengupta, Mark Ibrahim, and Laurens van der Maaten · 2021
Later among the works it cites.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander Rush · 2021
Later among the works it cites.
Cryptgpu: Fast privacy-preserving machine learning on the gpu
Sijun Tan, Brian Knott, Yuan Tian, and David J. Wu · 2021
Later among the works it cites.
THE-X: Privacy-preserving transformer inference with homomorphic encryption
Tianyu Chen, Hangbo Bao, Shaohan Huang, Li Dong, Binxing Jiao, Daxin Jiang, Haoyi Zhou, Jianxin Li, and Furu Wei · 2022
Later among the works it cites.
Spdz2k: Efficient mpc mod 2 for dishonest majority
Ronald Cramer, Ivan Damgård, Daniel E. Escudero, Peter Scholl, and Chaoping Xing · 2022
Later among the works it cites.
Iron: Private inference on transformers
Meng Hao, Hongwei Li, Hanxiao Chen, Pengzhi Xing, Guowen Xu, and Tianwei Zhang · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu · 2019
Cited alongside, same era.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Delphi: A cryptographic inference system for neural networks
Pratyush Mishra, Ryan Lehmkuhl, Akshayaram Srinivasan, Wenting Zheng, and Raluca Ada Popa · 2020
Cited alongside, same era.
Cryptflow2: Practical 2-party secure inference
Deevashwer Rathee, Mayank Rathee, Nishant Kumar, Nishanth Chandran, Divya Gupta, Aseem Rastogi, and Rahul Sharma · 2020
Cited alongside, same era.
Ferret: Fast extension for correlated ot with small communication
Kang Yang, Chenkai Weng, Xiao Lan, Jiang Zhang, and Xiao Wang · 2020
Cited alongside, same era.
Later among the works it cites.
Cheetah: Lean and fast secure Two-Party deep neural network inference
Zhicong Huang, Wen jie Lu, Cheng Hong, and Jiansheng Ding · 2022
Later among the works it cites.
Mpcformer: fast, performant and private transformer inference with mpc
Dacheng Li, Rulin Shao, Hongyi Wang, Han Guo, Eric P Xing, and Hao Zhang · 2022
Later among the works it cites.
Make Web3. 0 Connected
Zhuotao Liu, Yangxi Xiang, Jian Shi, Peng Gao, Haoyu Wang, Xusheng Xiao, Bihan Wen, Qi Li, and Yih-Chun Hu · 2022
Later among the works it cites.
Ppmlac: High performance chipset architecture for secure multi-party computation
Xing Zhou, Zhilei Xu, Cong Wang, and Mingyu Gao · 2022
Later among the works it cites.
martFL: Enabling Utility-Driven Data Marketplace with a Robust and Verifiable Federated Learning Architecture
Qi Li, Zhuotao Liu, Qi Li, and Ke Xu · 2023
Closest in time.
Efficient 3PC for binary circuits with application to Maliciously-Secure DNN inference
Yun Li, Yufei Duan, Zhicong Huang, Cheng Hong, Chao Zhang, and Yifan Song · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.