Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are not only resource-intensive to train but even more costly to deploy in production.
A value for n-person games
Shapley, L. S. et al · 1953
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
Veit, A., Wilber, M. J., and Belongie, S · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Convolutional networks with adaptive inference graphs
Veit, A. and Belongie, S · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Revisiting model stitching to compare neural representations
Bansal, Y., Nakkiran, P., and Barak, B · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
van der Ouderaa, T. F., Nagel, M., Van Baalen, M., Asano, Y. M., and Blankevoort, T · 2023
Later among the works it cites.
Efficient large language models: A survey
Wan, Z., Wang, X., Liu, C., Alam, S., Zheng, Y., Qu, Z., Yan, S., Zhu, Y., Zhang, Q., Chowdhury, M., et al · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
Xia, M., Gao, T., Zeng, Z., and Chen, D · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Toxigen: A large-scale machine-generated dataset for implicit and adversarial hate speech detection
Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., and Kamar, E · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Open llm leaderboard
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., Le Noac’h, A., Li, H., McDonell, K., Muennighoff, N., Ociepa, C., Phang, J., Reynolds, L., Schoelkopf, H., Skowron, A., Sutawika, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
https://github.com/jcpeterson/openwebtext , 2019
OpenWebText corpus · 2024
Closest in time.
Slicegpt: Compress large language models by deleting rows and columns
Ashkboos, S., Croci, M. L., Nascimento, M. G. d., Hoefler, T., and Hensman, J · 2024
Closest in time.
The unreasonable ineffectiveness of the deeper layers
Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., and Roberts, D. A · 2024
Closest in time.
Ffn-skipllm: A hidden gem for autoregressive decoding with adaptive feed forward skipping
Jaiswal, A., Hu, B., Yin, L., Ro, Y., Liu, S., Chen, T., and Akella, A · 2024
Closest in time.
Shortened llama: A simple depth pruning for large language models
Kim, B.-K., Kim, G., Kim, T.-H., Castells, T., Choi, S., Shin, J., and Song, H.-K · 2024
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect
Men, X., Xu, M., Zhang, Q., Wang, B., Lin, H., Lu, Y., Han, X., and Chen, W · 2024
Closest in time.
Dynamic memory compression: Retrofitting llms for accelerated inference
Nawrot, P., Łańcucki, A., Chochowski, M., Tarjan, D., and Ponti, E. M · 2024
Closest in time.
Mixture-of-depths: Dynamically allocating compute in transformer-based language models
Raposo, D., Ritter, S., Richards, B., Lillicrap, T., Humphreys, P. C., and Santoro, A · 2024
Closest in time.