Fetching the paper…
Reading the bibliography…
Rapid advancements in GPU computational power has outpaced memory capacity and bandwidth growth, creating bottlenecks in Large Language Model (LLM) inference.
Low-bit quantization of neural networks for efficient inference, 2019
Y. Choukroun, E. Kravchik, F. Yang, and P. Kisilev · 1902
Earlier work this paper cites.
A study of bfloat16 for deep learning training, 2019
D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. Vooturi, N. Jammalamadaka, J. Huang, H. Yuen, J. Yang, J. Park, A. Heinecke, E. Georganas, S. Srinivasan, A. Kundu, M. Smelyanskiy, B. Kaul, and P. Dubey · 1905
Earlier work this paper cites.
A method for the solution of certain non-linear problems in least squares
K. Levenberg · 1944
Earlier work this paper cites.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
An algorithm for least-squares estimation of nonlinear parameters
D. Marquardt · 1963
Earlier work this paper cites.
Zeroq: A novel zero shot quantization framework, 2020
Y. Cai, Z. Yao, Z. Dong, A. Gholami, M. W. Mahoney, and K. Keutzer · 2001
Earlier work this paper cites.
Scaling laws for neural language models, 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Probability, Random Variables, and Stochastic Processes
A. Papoulis and S. U. Pillai · 2002
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
D. J. C. MacKay · 2003
Earlier work this paper cites.
Language models are few-shot learners, 2020b
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
Improving post training neural quantization: Layer-wise calibration and integer programming, 2020
I. Hubara, Y. Nahshan, Y. Hanani, R. Banner, and D. Soudry · 2006
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation, 2013
Y. Bengio, N. Léonard, and A. Courville · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space, 2013
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
F. Li, B. Zhang, and B. Liu · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2016
Earlier work this paper cites.
Binarized neural networks
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
D. P. Kingma and J. Ba · 2017
Earlier work this paper cites.
Using the output embedding to improve language models, 2017
O. Press and L. Wolf · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
J. Welbl, N. F. Liu, and M. Gardner · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Earlier work this paper cites.
C. Zhu, S. Han, H. Mao, and W. J. Dally · 2017
Earlier work this paper cites.
Training competitive binary neural networks from scratch
J. Bethge, M. Bornstein, A. Loy, H. Yang, and C. Meinel · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Mixed precision training, 2018
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu · 2018
Earlier work this paper cites.
Towards understanding the role of over-parametrization in generalization of neural networks
B. Neyshabur, Z. Li, S. Bhojanapalli, Y. LeCun, and N. Srebro · 2018
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients, 2018
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. L. Bras, J. Gao, and Y. Choi · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2019
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Experiment tracking with weights and biases, 2020
L. Biewald · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Cited alongside, same era.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
N. Nangia, C. Vania, R. Bhalerao, and S. R. Bowman · 2020
Cited alongside, same era.
Zero: memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
Glu variants improve transformer, 2020
N. Shazeer · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Amd instinct mi210 accelerator
AMD Team · 2024
Closest in time.
Amd instinct mi250 and mi250x accelerators
AMD Team · 2024
Closest in time.
Amd instinct mi300a accelerator
AMD Team · 2024
Closest in time.
Amd instinct mi300x accelerator
AMD Team · 2024
Closest in time.
Amd instinct mi325x accelerator
AMD Team · 2024
Closest in time.
xlstm: Extended long short-term memory
M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush · 2020
Cited alongside, same era.
Understanding and overcoming the challenges of efficient transformer quantization, 2021
Y. Bondarenko, M. Nagel, and T. Blankevoort · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods, 2021
S. Lin, J. Hilton, and O. Evans · 2021
Cited alongside, same era.
Logiqa: a challenge dataset for machine reading comprehension with logical reasoning
J. Liu, L. Cui, H. Liu, D. Huang, Y. Wang, and Y. Zhang · 2021
Cited alongside, same era.
Winogrande: an adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2021
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu · 2021
Cited alongside, same era.
T. Dao and A. Gu · 2024
Closest in time.
Extreme compression of large language models via additive quantization, 2024
V. Egiazarian, A. Panferov, D. Kuznedelev, E. Frantar, A. Babenko, and D. Alistarh · 2024
Closest in time.
Marlin: a fast 4-bit inference kernel for medium batchsizes
E. Frantar and D. Alistarh · 2024
Closest in time.
A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer · 2024
Closest in time.
Google cloud tpu v3
Google TPU Team · 2024
Closest in time.
Google cloud tpu v4
Google TPU Team · 2024
Closest in time.
Google cloud tpu v5e
Google TPU Team · 2024
Closest in time.
Google cloud tpu v5p
Google TPU Team · 2024
Closest in time.
Olmo: Accelerating the science of language models, 2024
D. Groeneveld, I. Beltagy, P. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. H. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. R. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. E. Peters, V. Pyatkin, A. Ravichander, D. Schwenk, S. Shah, W. Smith, E. Strubell, N. Subramani, M. Wortsman, P. Dasigi, N. Lambert, K. Richardson, L. Zettlemoyer, J. Dodge, K. Lo, L. Soldaini, N. A. Smith, and H. Hajishirzi · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
A. Gu and T. Dao · 2024
Closest in time.
Minicpm: Unveiling the potential of small language models with scalable training strategies, 2024
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, X. Zhang, Z. L. Thai, K. Zhang, C. Wang, Y. Yao, C. Zhao, J. Zhou, J. Cai, Z. Zhai, N. Ding, C. Jia, G. Zeng, D. Li, Z. Liu, and M. Sun · 2024
Closest in time.
How good are low-bit quantized llama3 models? an empirical study, 2024
W. Huang, X. Ma, H. Qin, X. Zheng, C. Lv, H. Chen, J. Luo, X. Qi, X. Liu, and M. Magno · 2024
Closest in time.
Intel gaudi 2 and gaudi 3 ai accelerators
Intel Gaudi Team · 2024
Closest in time.
Squeezellm: Dense-and-sparse quantization, 2024
S. Kim, C. Hooper, A. Gholami, Z. Dong, X. Li, S. Shen, M. W. Mahoney, and K. Keutzer · 2024
Closest in time.
C. Lee, J. Jin, T. Kim, H. Kim, and E. Park · 2024
Closest in time.
Evaluating quantized large language models, 2024
S. Li, X. Ning, L. Wang, T. Liu, X. Shi, S. Yan, G. Dai, H. Yang, and Y. Wang · 2024
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration, 2024
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han · 2024
Closest in time.
The era of 1-bit llms: All large language models are in 1.58 bits, 2024
S. Ma, H. Wang, L. Ma, L. Wang, W. Wang, S. Huang, L. Dong, R. Wang, J. Xue, and F. Wei · 2024
Closest in time.
Pv-tuning: Beyond straight-through estimation for extreme llm compression, 2024
V. Malinovskii, D. Mazur, I. Ilin, D. Kuznedelev, K. Burlachenko, K. Yi, D. Alistarh, and P. Richtarik · 2024
Closest in time.
Nemotron-4 340b technical report, 2024
Nvidia, :, B. Adler, N. Agarwal, A. Aithal, D. H. Anh, P. Bhattacharya, A. Brundyn, J. Casper, B. Catanzaro, S. Clay, J. Cohen, S. Das, A. Dattagupta, O. Delalleau, L. Derczynski, Y. Dong, D. Egert, E. Evans, A. Ficek, D. Fridman, S. Ghosh, B. Ginsburg, I. Gitman, T. Grzegorzek, R. Hero, J. Huang, V. Jawa, J. Jennings, A. Jhunjhunwala, J. Kamalu, S. Khan, O. Kuchaiev, P. LeGresley, H. Li, J. Liu, Z. Liu, E. Long, A. S. Mahabaleshwarkar, S. Majumdar, J. Maki, M. Martinez, M. R. de Melo, I. Moshkov, D. Narayanan, S. Narenthiran, J. Navarro, P. Nguyen, O. Nitski, V. Noroozi, G. Nutheti, C. Parisien, J. Parmar, M. Patwary, K. Pawelec, W. Ping, S. Prabhumoye, R. Roy, T. Saar, V. R. N. Sabavat, S. Satheesh, J. P. Scowcroft, J. Sewall, P. Shamis, G. Shen, M. Shoeybi, D. Sizer, M. Smelyanskiy, F. Soares, M. N. Sreedhar, D. Su, S. Subramanian, S. Sun, S. Toshniwal, H. Wang, Z. Wang, J. You, J. Zeng, J. Zhang, J. Zhang, V. Zhang, Y. Zhang, and C. Zhu · 2024
Closest in time.
Nvidia tesla v100 gpu accelerator
Nvidia Team · 2024
Closest in time.
Nvidia a100 tensor core gpu
Nvidia Team · 2024
Closest in time.
Nvidia h100 tensor core gpu
Nvidia Team · 2024
Closest in time.
Nvidia h200 tensor core gpu
Nvidia Team · 2024
Closest in time.
Nvidia blackwell architecture
Nvidia Team · 2024
Closest in time.
Accelerating triton with torchscript and tensorrt integration
PyTorch Team · 2024
Closest in time.
Gemma: Open models based on gemini research and technology, 2024
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, P. G. Sessa, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. Héliou, A. Tacchetti, A. Bulanova, A. Paterson, B. Tsai, B. Shahriari, C. L. Lan, C. A. Choquette-Choo, C. Crepy, D. Cer, D. Ippolito, D. Reid, E. Buchatskaya, E. Ni, E. Noland, G. Yan, G. Tucker, G.-C. Muraru, G. Rozhdestvenskiy, H. Michalewski, I. Tenney, I. Grishchenko, J. Austin, J. Keeling, J. Labanowski, J.-B. Lespiau, J. Stanway, J. Brennan, J. Chen, J. Ferret, J. Chiu, J. Mao-Jones, K. Lee, K. Yu, K. Millican, L. L. Sjoesund, L. Lee, L. Dixon, M. Reid, M. Mikuła, M. Wirth, M. Sharman, N. Chinaev, N. Thain, O. Bachem, O. Chang, O. Wahltinez, P. Bailey, P. Michel, P. Yotov, R. Chaabouni, R. Comanescu, R. Jana, R. Anil, R. McIlroy, R. Liu, R. Mullins, S. L. Smith, S. Borgeaud, S. Girgin, S. Douglas, S. Pandya, S. Shakeri, S. De, T. Klimenko, T. Hennigan, V. Feinberg, W. Stokowiec, Y. hui Chen, Z. Ahmed, Z. Gong, T. Warkentin, L. Peran, M. Giang, C. Farabet, O. Vinyals, J. Dean, K. Kavukcuoglu, D. Hassabis, Z. Ghahramani, D. Eck, J. Barral, F. Pereira, E. Collins, A. Joulin, N. Fiedel, E. Senter, A. Andreev, and K. Kenealy · 2024
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models, 2024
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2024
Closest in time.
PB-LLM: Partially binarized large language models
Z. Yuan, Y. Shang, and Z. Dong · 2024
Closest in time.