Fetching the paper…
Reading the bibliography…
Privacy-preserving computation enables language model inference directly on encrypted data yet suffers from prohibitive latency and communication overheads, primarily due to nonlinear functions.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Information theory and statistical mechanics
Jaynes, E. T · 1957
Earlier work this paper cites.
Perplexity—a measure of the difficulty of speech recognition tasks
Jelinek, F., Mercer, R. L., Bahl, L. R., and Baker, J. K · 1977
Earlier work this paper cites.
On the rationale of maximum-entropy methods
Jaynes, E. T · 1982
Earlier work this paper cites.
A global optimization technique for statistical classifier design
Miller, D., Rao, A. V., Rose, K., and Gersho, A · 1996
Earlier work this paper cites.
Extending oblivious transfers efficiently
Ishai, Y., Kilian, J., Nissim, K., and Petrank, E · 2003
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
What does attention in neural machine translation pay attention to?
Ghader, H. and Monz, C · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Earlier work this paper cites.
Regularizing neural networks by penalizing confident output distributions
Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., and Hinton, G · 2017
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Earlier work this paper cites.
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D · 2019
Earlier work this paper cites.
Information-theoretic confidence bounds for reinforcement learning
Lu, X. and Van Roy, B · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Vig, J. and Belinkov, Y · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Earlier work this paper cites.
Roles and utilization of attention heads in transformer-based neural language models
Jo, J.-y. and Myaeng, S.-H · 2020
Earlier work this paper cites.
Compositional explanations of neurons
Mu, J. and Andreas, J · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Earlier work this paper cites.
Ferret: Fast extension for correlated ot with small communication
Yang, K., Weng, C., Lan, X., Zhang, J., and Wang, X · 2020
Earlier work this paper cites.
Domain generalization via entropy regularization
Zhao, S., Gong, M., Liu, T., Fu, H., and Tao, D · 2020
Earlier work this paper cites.
Crypten: Secure multi-party computation meets machine learning
Knott, B., Venkataraman, S., Hannun, A., Sengupta, S., Ibrahim, M., and van der Maaten, L · 2021
Earlier work this paper cites.
Bert busters: Outlier dimensions that disrupt transformers
Kovaleva, O., Kulshreshtha, S., Rogers, A., and Rumshisky, A · 2021
Earlier work this paper cites.
Contributions of transformer attention heads in multi- and cross-lingual tasks
Ma, W., Zhang, K., Lou, R., Wang, L., and Vosoughi, S · 2021
Earlier work this paper cites.
Iron: Private inference on transformers
Hao, M., Li, H., Chen, H., Xing, P., Xu, G., and Zhang, T · 2022
Earlier work this paper cites.
Block-recurrent transformers
Hutchins, D., Schlag, I., Wu, Y., Dyer, E., and Neyshabur, B · 2022
Earlier work this paper cites.
Adversarially robust learning via entropic regularization
Jagatap, G., Joshi, A., Chowdhury, A. B., Garg, S., and Hegde, C · 2022
Earlier work this paper cites.
Accelerating attention through gradient-based learned runtime pruning
Li, Z., Ghodrati, S., Yazdanbakhsh, A., Esmaeilzadeh, H., and Kang, M · 2022
Earlier work this paper cites.
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Noci, L., Anagnostidis, S., Biggio, L., Orvieto, A., Singh, S. P., and Lucchi, A · 2022
Earlier work this paper cites.
Improving the trainability of deep neural networks through layerwise batch-entropy regularization
Peer, D., Keulen, B., Stabinger, S., Piater, J., and Rodriguez-sanchez, A · 2022
Earlier work this paper cites.
Maximizing entropy on adversarial examples can improve generalization
Setlur, A., Eysenbach, B., Smith, V., and Levine, S · 2022
Earlier work this paper cites.
Outlier suppression: Pushing the limit of low-bit transformer language models
Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X · 2022
Cited alongside, same era.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2023
Cited alongside, same era.
On the expressivity role of layernorm in transformers’ attention
Brody, S., Alon, U., and Yahav, E · 2023
Cited alongside, same era.
Rna-vit: Reduced-dimension approximate normalized attention vision transformers for latency efficient private inference
Chen, D., Zhang, Y., Kundu, S., Li, C., and Beerel, P. A · 2023
Cited alongside, same era.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Cited alongside, same era.
Understanding and minimising outlier features in neural network training
He, B., Noci, L., Paliotta, D., Schlag, I., and Hofmann, T · 2024
Closest in time.
Can perplexity reflect large language model’s ability in long text understanding?
Hu, Y., Huang, Q., Tao, M., Zhang, C., and Feng, Y · 2024
Closest in time.
Distillm: Towards streamlined distillation for large language models
Ko, J., Kim, S., Chen, T., and Yun, S.-Y · 2024
Closest in time.
How do nonlinear transformers learn and generalize in in-context learning?
Li, H., Wang, M., Lu, S., Cui, X., and Chen, P.-Y · 2024
Closest in time.
How does architecture influence the base capabilities of pre-trained language models? a case study based on ffn-wider transformer models
Lu, X., Zhao, Y., and Qin, B · 2024
Closest in time.
Can LLMs keep a secret? testing privacy implications of language models via contextual integrity theory
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dettmers, T. and Zettlemoyer, L · 2023
Cited alongside, same era.
Cramming: Training a language model on a single gpu in one day
Geiping, J. and Goldstein, T · 2023
Cited alongside, same era.
Minillm: Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Cited alongside, same era.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
He, B., Martens, J., Zhang, G., Botev, A., Brock, A., Smith, S. L., and Teh, Y. W · 2023
Cited alongside, same era.
Ciphergpt: Secure two-party gpt inference
Hou, X., Liu, J., Li, J., Li, Y., Lu, W.-j., Hong, C., and Ren, K · 2023
Cited alongside, same era.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Cited alongside, same era.
Mireshghallah, N., Kim, H., Zhou, X., Tsvetkov, Y., Sap, M., Shokri, R., and Choi, Y · 2024
Closest in time.
ReLU strikes back: Exploiting activation sparsity in large language models
Mirzadeh, S. I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C. C., Tuzel, O., Samei, G., Rastegari, M., and Farajtabar, M · 2024
Closest in time.
Olmoe: Open mixture-of-experts language models
Muennighoff, N., Soldaini, L., Groeneveld, D., Lo, K., Morrison, J., Min, S., Shi, W., Walsh, P., Tafjord, O., Lambert, N., et al · 2024
Closest in time.
Linear log-normal attention with unbiased concentration
Nahshan, Y., Kampeas, J., and Haleva, E · 2024
Closest in time.
On the nonlinearity of layer normalization
Ni, Y., Guo, Y., Jia, J., and Huang, L · 2024
Closest in time.
Bolt: Privacy-preserving, accurate and efficient inference for transformers
Pang, Q., Zhu, J., Möllering, H., Zheng, W., and Schneider, T · 2024
Closest in time.
Weight-based decomposition: A case for bilinear MLPs
Pearce, M. T., Dooms, T., and Rigg, A · 2024
Closest in time.
Methods of improving llm training stability
Rybakov, O., Chrzanowski, M., Dykas, P., Xue, J., and Lanir, B · 2024
Closest in time.
Beyond memorization: Violating privacy via inference with large language models
Staab, R., Vero, M., Balunovic, M., and Vechev, M · 2024
Closest in time.
Diffusion actor-critic with entropy regulator
Wang, Y., Wang, L., Jiang, Y., Zou, W., Liu, T., Song, X., Wang, W., Xiao, L., Wu, J., Duan, J., et al · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
Wortsman, M., Liu, P. J., Xiao, L., Everett, K. E., Alemi, A. A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., Pennington, J., Sohl-Dickstein, J., Xu, K., Lee, J., Gilmer, J., and Kornblith, S · 2024
Closest in time.
On the role of attention masks and layernorm in transformers
Wu, X., Ajorlou, A., Wang, Y., Jegelka, S., and Jadbabaie, A · 2024
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Closest in time.
Ye, T., Dong, L., Xia, Y., Sun, Y., Zhu, Y., Huang, G., and Wei, F · 2024
Closest in time.
Stablemask: Refining causal masking in decoder-only transformer
Yin, Q., He, X., Zhuang, X., Zhao, Y., Yao, J., Shen, X., and Zhang, Q · 2024
Closest in time.
Unveiling and harnessing hidden attention sinks: Enhancing large language models without training through attention calibration
Yu, Z., Wang, Z., Fu, Y., Shi, H., Shaikh, K., and Lin, Y. C · 2024
Closest in time.
The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry
Zhang, M., Bhatia, K., Kumbong, H., and Re, C · 2024
Closest in time.
Tuning LayerNorm in attention: Towards efficient multi-modal llm finetuning
Zhao, B., Tu, H., Wei, C., Mei, J., and Xie, C · 2024
Closest in time.
Converting transformers to polynomial form for secure inference over homomorphic encryption
Zimerman, I., Baruch, M., Drucker, N., Ezov, G., Soceanu, O., and Wolf, L · 2024
Closest in time.
What is wrong with perplexity for long-context language modeling?
Fang, L., Wang, Y., Liu, Z., Zhang, C., Jegelka, S., Gao, J., Ding, B., and Wang, Y · 2025
Closest in time.
Mitigating attention localization in small scale: Self-attention refinement via one-step belief propagation
Lee, N., Kim, Y., Oh, M., Kim, S., Koo, J. W., Jo, H., and Lee, J · 2025
Closest in time.
Mix-LN: Unleashing the power of deeper layers by combining pre-LN and post-LN
Li, P., Yin, L., and Liu, S · 2025
Closest in time.
Bumblebee: Secure two-party inference framework for large transformers
Lu, W.-j., Huang, Z., Gu, Z., Li, J., Liu, J., Ren, K., Hong, C., Wei, T., and Chen, W · 2025
Closest in time.
Thor: Secure transformer inference with homomorphic encryption
Moon, J., Yoo, D., Jiang, X., and Kim, M · 2025
Closest in time.
Softmax is not enough (for sharp size generalisation)
Veličković, P., Perivolaropoulos, C., Barbero, F., and Pascanu, R · 2025
Closest in time.
Breaking the layer barrier: Remodeling private transformer inference with hybrid { \{ CKKS } \} and { \{ MPC } \}
Xu, T., Lu, W.-j., Yu, J., Chen, Y., Lin, C., Wang, R., and Li, M · 2025
Closest in time.
Transformers without normalization
Zhu, J., Chen, X., He, K., LeCun, Y., and Liu, Z · 2025
Closest in time.