Fetching the paper…
Reading the bibliography…
Monolithic large language models (LLMs) like GPT-4 have paved the way for modern generative AI applications.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural Computation , p. 79–87, Jan 1991. [Online]. Available: http://dx.doi.org/10.1162/neco.1991.3.1.79
1991
Earlier work this paper cites.
S. Williams, A. Waterman, and D. Patterson, “Roofline: An insightful visual performance model for multicore architectures,” Communications of the ACM , vol. 52, no. 4, pp. 65–76, 2009. [Online]. Available: https://dl.acm.org/doi/10.1145/1498765.1498785
2009
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa et al. , “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the ACM/IEEE International Symposium on Computer Architecture (ISCA) , 2017
2017
Earlier work this paper cites.
L. Liu, J. Zhu, Z. Li, Y. Lu, Y. Deng, J. Han et al. , “A survey of coarse-grained reconfigurable architecture and design: Taxonomy, challenges, and applications,” ACM Comput. Surv. , vol. 52, no. 6, oct 2019. [Online]. Available: https://doi.org/10.1145/3357375
2019
Earlier work this paper cites.
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon et al. , “A domain-specific supercomputer for training deep neural networks,” Communications of the ACM , vol. 63, no. 7, pp. 67–78, Jun. 2020
2020
Earlier work this paper cites.
A. Podobas, K. Sano, and S. Matsuoka, “A survey on coarse-grained reconfigurable architectures from a performance perspective,” IEEE Access , vol. 8, p. 146719–146743, 2020. [Online]. Available: http://dx.doi.org/10.1109/ACCESS.2020.3012084
2020
Earlier work this paper cites.
J. Choquette, W. Gandhi, O. Giroux, N. Stam, and R. Krashinsky, “Nvidia a100 tensor core gpu: Performance and innovation,” IEEE Micro , vol. 41, no. 2, pp. 29–35, 2021
2021
Earlier work this paper cites.
N. P. Jouppi, D. H. Yoon, M. Ashcraft, M. Gottscho, T. B. Jablin, G. Kurian et al. , “Ten lessons from three generations shaped google’s tpuv4i: Industrial product,” in Proceedings of the ACM/IEEE International Symposium on Computer Architecture (ISCA) , 2021
2021
Earlier work this paper cites.
A. Ivanov, N. Dryden, T. Ben-Nun, S. Li, and T. Hoefler, “Data movement is all you need: A case study on optimizing transformers,” 2021
2021
Earlier work this paper cites.
R. Prabhakar and S. Jairath, “Sambanova sn10 rdu:accelerating software 2.0 with dataflow,” in 2021 IEEE Hot Chips 33 Symposium (HCS) , 2021, pp. 1–37
2021
Earlier work this paper cites.
S. Knowles, “Graphcore,” in 2021 IEEE Hot Chips 33 Symposium (HCS) , 2021, pp. 1–25
2021
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang et al. , “Lora: Low-rank adaptation of large language models,” 2021
2021
Earlier work this paper cites.
J. Choquette, “Nvidia hopper gpu: Scaling performance,” in 2022 IEEE Hot Chips 34 Symposium (HCS) , 2022, pp. 1–46
2022
Earlier work this paper cites.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus et al. , “Scaling instruction-finetuned language models,” 2022. [Online]. Available: https://doi.org/10.48550/arXiv.2210.11416
2022
Earlier work this paper cites.
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai et al. , “Multitask prompted training enables zero-shot task generalization,” 2022. [Online]. Available: https://doi.org/10.48550/arXiv.2110.08207
2022
Earlier work this paper cites.
Y. Wang, S. Mishra, P. Alipoormolabashi, Y. Kordi, A. Mirzaei, A. Arunkumar et al. , “Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks,” 2022. [Online]. Available: https://doi.org/10.48550/arXiv.2204.07705
2022
Earlier work this paper cites.
M. Li, S. Gururangan, T. Dettmers, M. Lewis, T. Althoff, N. A. Smith et al. , “Branch-train-merge: Embarrassingly parallel training of expert language models,” 2022
2022
Earlier work this paper cites.
D. Zhang, S. Huda, E. Songhori, K. Prabhu, Q. Le, A. Goldie et al. , “A full-stack search technique for domain optimized deep learning accelerators,” in Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 27–42. [Online]. Available: https://doi.org/10.1145/3503222.3507767
2022
Earlier work this paper cites.
R. Prabhakar, S. Jairath, and J. L. Shin, “Sambanova sn10 rdu: A 7nm dataflow architecture to accelerate software 2.0,” in 2022 IEEE International Solid-State Circuits Conference (ISSCC) , vol. 65, 2022, pp. 350–352
2022
Cited alongside, same era.
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “Flashattention: Fast and memory-efficient exact attention with io-awareness,” 2022
2022
Cited alongside, same era.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya et al. , “Efficiently scaling transformer inference,” 2022
2022
Cited alongside, same era.
S. Lie, “Cerebras architecture deep dive: First look inside the hw/sw co-design for deep learning : Cerebras systems,” in 2022 IEEE Hot Chips 34 Symposium (HCS) , 2022, pp. 1–34
2022
Cited alongside, same era.
T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning,” 2023
2023
Later among the works it cites.
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” 2023
2023
Later among the works it cites.
B. Workshop, :, T. L. Scao, A. Fan, C. Akiki, E. Pavlick et al. , “Bloom: A 176b-parameter open-access multilingual language model,” 2023
2023
Later among the works it cites.
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Debbah et al. , “The falcon series of open language models,” 2023
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Abts, G. Kimmell, A. Ling, J. Kim, M. Boyd, A. Bitar et al. , “A software-defined tensor streaming multiprocessor for large-scale machine learning,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , ser. ISCA ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 567–580. [Online]. Available: https://doi.org/10.1145/3470496.3527405
2022
Cited alongside, same era.
N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai et al. , “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” in Proceedings of the ACM/IEEE International Symposium on Computer Architecture (ISCA) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei et al. , “Llama 2: Open foundation and fine-tuned chat models,” 2023. [Online]. Available: https://doi.org/10.48550/arXiv.2307.09288
2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas et al. , “Mistral 7b,” 2023
2023
Cited alongside, same era.
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan et al. , “Code llama: Open foundation models for code,” 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
H. Ivison, Y. Wang, V. Pyatkin, N. Lambert, M. Peters, P. Dasigi et al. , “Camels in a changing climate: Enhancing lm adaptation with tulu 2,” 2023
2023
Later among the works it cites.
W. Zou, Q. Li, J. Ge, C. Li, X. Shen, L. Huang et al. , “A comprehensive evaluation of parameter-efficient fine-tuning on software engineering tasks,” 2023
2023
Later among the works it cites.
N. Maslej, L. Fattorini, R. Perrault, V. Parli, A. Reuel, E. Brynjolfsson et al. , ““the ai index 2024 annual report,” ai index steering committee, institute for human-centered ai, stanford university,” 2024. [Online]. Available: https://aiindex.stanford.edu/wp-content/uploads/2024/04/HAI_AI-Index-Report-2024.pdf
2024
Closest in time.
A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, and K. Keutzer, “Ai and memory wall,” 2024
2024
Closest in time.
2024
Closest in time.
X. Lu, A. Liusie, V. Raina, Y. Zhang, and W. Beauchamp, “Blending is all you need: Cheaper, better alternative to trillion-parameters llm,” 2024
2024
Closest in time.
M. Zaharia, O. Khattab, L. Chen, J. Q. Davis, H. Miller, C. Potts et al. , “The shift from models to compound ai systems,” https://bair.berkeley.edu/blog/2024/02/18/compound-ai-systems/ , 2024
2024
Closest in time.
A. Ng, “The batch, issue 246,” https://www.deeplearning.ai/the-batch/issue-246/ , 2024
2024
Closest in time.
A. Analysis, “Llama 3.1 405b: Api provider benchmarking and analysis,” https://artificialanalysis.ai/models/llama-3-1-instruct-405b/providers ,, September 2024
2024
Closest in time.
N. Mundra, S. Doddapaneni, R. Dabre, A. Kunchukuttan, R. Puduppully, and M. M. Khapra, “A comprehensive analysis of adapter efficiency,” in Proceedings of the 7th Joint International Conference on Data Science & Management of Data (11th ACM IKDD CODS and 29th COMAD) , ser. CODS-COMAD ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 136–154. [Online]. Available: https://doi.org/10.1145/3632410.3632463
2024
Closest in time.
B. Zhang, Z. Liu, C. Cherry, and O. Firat, “When scaling meets llm finetuning: The effect of data, model and finetuning method,” 2024
2024
Closest in time.