Fetching the paper…
Reading the bibliography…
This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton · 2015
Earlier work this paper cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Raghuraman Krishnamoorthi · 2018
Earlier work this paper cites.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart Van Baalen, and Tijmen Blankevoort · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Privileged bases in the transformer residual stream, 2023
Nelson Elhage, Roberb Lasenby, and Christopher Olah · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa · 2023
Cited alongside, same era.
Llm-qat: Data-free quantization aware training for large language models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra · 2023
Cited alongside, same era.
https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 2024 · 2024
Cited alongside, same era.
The llama 3 family of models
AI @ Meta Llama Team · 2024
Cited alongside, same era.
Shortgpt: Layers in large language models are more redundant than you expect, 2024
Announcing mlcommons ai safety v0.5 proof of concept
MLCommons · 2024
Closest in time.
Compact language models via pruning and knowledge distillation, 2024
Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, and Pavlo Molchanov · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al · 2024
Closest in time.
torchao: Pytorch native quantization and sparsity for training and inference, October 2024
torchao maintainers and contributors · 2024
Closest in time.
torchtune: Pytorch’s finetuning library, April 2024
torchtune maintainers and contributors · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xin Men, Mingyu Xu, Qingyu Zhang, Bingning Wang, Hongyu Lin, Yaojie Lu, Xianpei Han, and Weipeng Chen · 2024
Cited alongside, same era.
Executorch llama android demo app
Executorch Team
Cited in the paper.
Executorch llama ios demo app
Executorch Team
Cited in the paper.
The llama 3 herd of models, 2024a
Llama Team
Cited in the paper.
Meta llama guard 2
Llama Team
Cited in the paper.
Executorch runtime overview
Pytorch Team
Cited in the paper.
Executorch xnnpack delegate
Pytorch Team
Cited in the paper.
Closest in time.