Fetching the paper…
Reading the bibliography…
Many companies use large language models (LLMs) offered as a service, like OpenAI's GPT-4, to create AI-enabled product experiences.
1908
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
C. D. Manning, An introduction to information retrieval . Cambridge university press, 2009
2009
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford and at el., “Language models are unsupervised multitask learners,” 2019. [Online]. Available: https://api.semanticscholar.org/CorpusID:160025533
2019
Earlier work this paper cites.
M. Shoeybi and at el., “Megatron-lm: Training multi-billion parameter language models using model parallelism,” 2020
2020
Earlier work this paper cites.
H. Kwon, L. Lai, M. Pellauer, T. Krishna, Y.-H. Chen, and V. Chandra, “Heterogeneous dataflow accelerators for multi-dnn workloads,” in 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2021, pp. 71–83
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
“Emnlp: Prompt engineering is the new feature engineering,” 2022. [Online]. Available: https://www.amazon.science/blog/emnlp-prompt-engineering-is-the-new-feature-engineering
2022
Earlier work this paper cites.
F. Blanaru, A. Stratikopoulos, J. Fumero, and C. Kotselidis, “Enabling pipeline parallelism in heterogeneous managed runtime environments via batch processing,” in Proceedings of the 18th ACM SIGPLAN/SIGOPS International Conference on Virtual Execution Environments , 2022, pp. 58–71
2022
Earlier work this paper cites.
A. Yazdanbakhsh and et al., “Sparse attention acceleration with synergistic in-memory pruning and on-chip recomputation,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . Los Alamitos, CA, USA: IEEE Computer Society, oct 2022, pp. 744–762. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/MICRO56248.2022.00059
2022
Earlier work this paper cites.
S. Wang and at el., “Overlap communication with dependent computation via decomposition in large deep learning models,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 , ser. ASPLOS 2023. Association for Computing Machinery, 2022, p. 93–106
2022
Earlier work this paper cites.
S. Hong and at el., “Dfx: A low-latency multi-fpga appliance for accelerating transformer-based text generation,” in 2022 IEEE Hot Chips 34 Symposium (HCS) , 2022, pp. 1–17
2022
Earlier work this paper cites.
“OpenAI blog: ChatGPT,” https://openai.com/blog/chatgpt , 2023
2023
Earlier work this paper cites.
M. Dowling and B. Lucey, “Chatgpt for (finance) research: The bananarama conjecture,” Finance Research Letters , vol. 53, p. 103662, 2023
2023
Earlier work this paper cites.
The Register. (2023) Outage hits OpenAI’s ChatGPT as servers stopped by ’Claude’ glitch. [Online]. Available: https://www.theregister.com/2023/11/08/outage˙chatgpt˙openai˙claude/
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7b,” 2023
2023
Earlier work this paper cites.
“neuralchat model listing,” https://huggingface.co/Intel/neural-chat-7b-v1-1 , 2023
2023
Earlier work this paper cites.
G. Wang, S. Cheng, Q. Yu, and C. Liu, “OpenLLMs: Less is More for Open-source Models,” 7 2023. [Online]. Available: https://github.com/imoneoi/openchat
2023
Earlier work this paper cites.
A. Mitra, L. D. Corro, S. Mahajan, A. Codas, C. Simoes, S. Agarwal, X. Chen, A. Razdaibiedina, E. Jones, K. Aggarwal, H. Palangi, G. Zheng, C. Rosset, H. Khanpour, and A. Awadallah, “Orca 2: Teaching small language models how to reason,” 2023
2023
Cited alongside, same era.
“Stabilitai: Stablelm,” https://github.com/Stability-AI/StableLM , 2023
2023
Cited alongside, same era.
B. Zhu, E. Frick, T. Wu, H. Zhu, and J. Jiao, “Starling-7b: Improving llm helpfulness and harmlessness with rlaif,” November 2023
2023
Cited alongside, same era.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023. [Online]. Available: https://lmsys.org/blog/2023-03-30-vicuna/
2023
Cited alongside, same era.
HuggingFace, “Hugging Face Models,” https://huggingface.co/models , 2023, accessed: Dec. 15, 2023
2023
Closest in time.
HuggingFaceH4, “Open LLM Leaderboard,” https://huggingface.co/spaces/HuggingFaceH4/open˙llm˙leaderboard , 2023, accessed: Dec. 15, 2023
2023
Closest in time.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging llm-as-a-judge with mt-bench and chatbot arena,” 2023
2023
Closest in time.
W. Won, T. Heo, S. Rashidi, S. Sridharan, S. Srinivasan, and T. Krishna, “Astra-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale,” in IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2023, Raleigh, NC, USA, April 23-25, 2023 . IEEE, 2023, pp. 283–294. [Online]. Available: https://doi.org/10.1109/ISPASS57527.2023.00035
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y. Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib, N. Sarrazin, O. Sanseviero, A. M. Rush, and T. Wolf, “Zephyr: Direct distillation of lm alignment,” 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Chang and at el., “A survey on evaluation of large language models,” 2023
2023
Cited alongside, same era.
M. Ollivier and at el., “A deeper dive into chatgpt: history, use and future perspectives for orthopaedic research,” Knee Surgery, Sports Traumatology, Arthroscopy: Official Journal of the ESSKA , vol. 31, no. 4, pp. 1190–1192, 2023
2023
Cited alongside, same era.
A. Vaswani and at el., “Attention is all you need,” 2023
2023
Cited alongside, same era.
D. Moolchandani, J. Kundu, F. Ruelens, P. Vrancx, T. Evenblij, and M. Perumkunnil, “Amped: An analytical model for performance in distributed training of transformers,” in IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2023, Raleigh, NC, USA, April 23-25, 2023 . IEEE, 2023, pp. 306–315. [Online]. Available: https://doi.org/10.1109/ISPASS57527.2023.00037
2023
Closest in time.
J. Gómez-Luna, Y. Guo, S. Brocard, J. Legriel, R. Cimadomo, G. F. Oliveira, G. Singh, and O. Mutlu, “Evaluating machine learningworkloads on memory-centric computing systems,” in 2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2023, pp. 35–49
2023
Closest in time.
H. Kwon, K. Nair, J. Seo, J. Yik, D. Mohapatra, D. Zhan, J. Song, P. Capak, P. Zhang, P. Vajda et al. , “Xrbench: An extended reality (xr) machine learning benchmark suite for the metaverse,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Wu, L. Y. Dai, A. Novick, M. Glick, Z. Zhu, S. Rumley, G. Michelogiannakis, J. Shalf, and K. Bergman, “Peta-scale embedded photonics architecture for distributed deep learning applications,” Journal of Lightwave Technology , 2023
2023
Closest in time.
J. Song, J. Yim, J. Jung, H. Jang, H.-J. Kim, Y. Kim, and J. Lee, “Optimus-cc: Efficient large nlp model training with 3d parallelism aware communication compression,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 560–573
2023
Closest in time.
S.-C. Kao, S. Subramanian, G. Agrawal, A. Yazdanbakhsh, and T. Krishna, “Flat: An optimized dataflow for mitigating attention bottlenecks,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 295–310
2023
Closest in time.
D. Rouhani and et al., “With shared microexponents, a little shifting goes a long way,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , ser. ISCA ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3579371.3589351
2023
Closest in time.
H. Zhang, A. Ning, R. Prabhakar, and D. Wentzlaff, “A hardware evaluation framework for large language model inference,” 2023
2023
Closest in time.
G. Jeong, S. Damani, A. R. Bambhaniya, E. Qin, C. J. Hughes, S. Subramoney, H. Kim, and T. Krishna, “Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 259–272
2023
Closest in time.
N. Jouppi and et al., “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , ser. ISCA ’23. Association for Computing Machinery, 2023
2023
Closest in time.
C. Guo and et al., “Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , ser. ISCA ’23. ACM, Jun. 2023. [Online]. Available: http://dx.doi.org/10.1145/3579371.3589038
2023
Closest in time.
Llamaindex, “Llamaindex, data framework for llm applications,” https://www.llamaindex.ai/ , 2024, accessed: April 13th, 2024
2024
Closest in time.
ollama, “ollama,” https://github.com/ollama/ollama , 2024, accessed: April 13th, 2024
2024
Closest in time.