Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) often excel in specific domains but fall short in others due to the limitations of their training.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Federated learning with matched averaging
Wang, H., Yurochkin, M., Sun, Y., Papailiopoulos, D., and Khazaeni, Y · 2020
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Git re-basin: Merging models modulo permutation symmetries
Ainsworth, S. K., Hayase, J., and Srinivasa, S · 2022
Earlier work this paper cites.
Hugging face
Jain, S. M · 2022
Earlier work this paper cites.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Rest: Retrieval-based speculative decoding
He, Z., Zhong, Z., Cai, T., Lee, J. D., and He, D · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Solar 10.7 b: Scaling large language models with simple yet effective depth up-scaling
Kim, D., Park, C., Kim, S., Lee, W., Song, W., Kim, Y., Kim, H., Kim, Y., Lee, H., Kim, J., et al · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Cited alongside, same era.
Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Luo, H., Sun, Q., Xu, C., Zhao, P., Lou, J., Tao, C., Geng, X., Lin, Q., Chen, S., and Zhang, D · 2023
Cited alongside, same era.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y., et al · 2024
Later among the works it cites.
Analysis of linear mode connectivity via permutation-based weight matching
Ito, A., Yamada, M., and Kumagai, A · 2024
Later among the works it cites.
Nearest neighbor speculative decoding for llm generation and attribution
Li, M., Chen, X., Holtzman, A., Chen, B., Lin, J., Yih, W.-t., and Lin, X. V · 2024
Later among the works it cites.
A novel committee-based framework for modeling groundwater level fluctuations: A combination of mathematical and machine learning models using the weighted multi-model ensemble mean algorithm
Mazraeh, A., Bagherifar, M., Shabanlou, S., and Ekhlasmand, R · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, H., Polo, F. M., Sun, Y., Kundu, S., Xing, E., and Yurochkin, M · 2023
Cited alongside, same era.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Xia, H., Ge, T., Wang, P., Chen, S.-Q., Wei, F., and Sui, Z · 2023
Cited alongside, same era.
Predictive pipelined decoding: A compute-latency trade-off for exact llm decoding
Yang, S., Lee, G., Cho, J., Papailiopoulos, D., and Lee, K · 2023
Cited alongside, same era.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S · 2023
Cited alongside, same era.
Distillspec: Improving speculative decoding via knowledge distillation
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2023
Cited alongside, same era.
Network intrusion detection system by applying ensemble model for smart home
Amru, M., Kannan, R. J., Ganesh, E. N., Muthumarilakshmi, S., Padmanaban, K., Jeyapriya, J., and Murugan, S · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Layer skip: Enabling early exit inference and self-speculative decoding
Elhoushi, M., Shrivastava, A., Liskovich, D., Hosmer, B., Wasti, B., Lai, L., Mahmoud, A., Acun, B., Agarwal, S., Roman, A., et al · 2024
Cited alongside, same era.
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Later among the works it cites.
tinybenchmarks: evaluating llms with fewer examples
Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M · 2024
Later among the works it cites.
Ensemble approach of transfer learning and vision transformer leveraging explainable ai for disease diagnosis: An advancement towards smart healthcare 5.0
Poonia, R. C. and Al-Alshaikh, H. A · 2024
Later among the works it cites.
Learning to decode collaboratively with multiple language models
Shen, S. Z., Lang, H., Wang, B., Kim, Y., and Sontag, D · 2024
Later among the works it cites.
Knowledge fusion of large language models
Wan, F., Huang, X., Cai, D., Quan, X., Bi, W., and Shi, S · 2024
Later among the works it cites.
Llama pro: Progressive llama with block expansion
Wu, C., Gan, Y., Ge, Y., Lu, Z., Wang, J., Feng, Y., Luo, P., and Shan, Y · 2024
Later among the works it cites.
Xia, H., Yang, Z., Dong, Q., Wang, P., Li, Y., Ge, T., Liu, T., Li, W., and Sui, Z · 2024
Later among the works it cites.
Openmoe: An early effort on open mixture-of-experts language models
Xue, F., Zheng, Z., Fu, Y., Ni, J., Zheng, Z., Zhou, W., and You, Y · 2024
Later among the works it cites.
Tinyllama: An open-source small language model
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Later among the works it cites.
Llama-moe: Building mixture-of-experts from llama with continual pre-training
Zhu, T., Qu, X., Dong, D., Ruan, J., Tong, J., He, C., and Cheng, Y · 2024
Later among the works it cites.