Fetching the paper…
Reading the bibliography…
The widespread adoption of large language models (LLMs) has created a pressing need for an efficient, secure and private serving infrastructure, which allows researchers to run open source or custom fine-tuned LLMs and ensures users that their data remains private and is not stored without their consent.
“Gradio: Hassle-free Sharing and Testing of ML Models in the Wild”, 2019
Abubakar Abid et al · 1906
Earlier work this paper cites.
“SLURM: Simple Linux Utility for Resource Management”
Andy. Yoo, Morris. Jette and Mark Grondona · 2003
Earlier work this paper cites.
“Introduction of a Web Service for Cloud Computing with the Integrated Hydrologic Simulation Platform ParFlow”
Claudius. Bürger, Stefan Kollet, Jens Schumacher and Detlef Bösel · 2012
Earlier work this paper cites.
“Introduction of a web service for cloud computing with the integrated hydrologic simulation platform ParFlow”
Claudius. Bürger, Stefan Kollet, Jens Schumacher and Detlef Bösel · 2012
Earlier work this paper cites.
“Addressing Security Aspects for HPC Infrastructure”
Ramesh Bulusu et al · 2018
Earlier work this paper cites.
“Deploying AI Frameworks on Secure HPC Systems with Containers”
David Brayford et al · 2019
Earlier work this paper cites.
“Artificial Intelligence Platform for Mobile Service Computing”
Haikuo Zhang et al · 2019
Earlier work this paper cites.
“Language Models Are Few-Shot Learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“Large-Scale HPC Deployment of Scalable CyberInfrastructure for Artificial Intelligence and Likelihood Free Inference (SCAILFIN)”
Michael Hildreth et al · 2020
Earlier work this paper cites.
“GPU-Accelerated Machine Learning Inference as a Service for Computing in Neutrino Experiments”
Michael Wang et al · 2020
Earlier work this paper cites.
“Container Orchestration on HPC Systems through Kubernetes”
Naweiluo Zhou et al · 2021
Earlier work this paper cites.
“PaLM: Scaling Language Modeling with Pathways”, 2022
Aakanksha Chowdhery et al · 2022
Earlier work this paper cites.
“FP8 Formats for Deep Learning”, 2022
Paulius Micikevicius et al · 2022
Earlier work this paper cites.
“NVIDIA’s Cloud Native Supercomputing”
Gilad Shainer et al · 2022
Earlier work this paper cites.
URL: https://huggingface.co/Intel/neural-chat-7b-v3-1
“Intel/Neural-Chat-7b-v3-1 · Hugging Face”, 2023 · 2023
Earlier work this paper cites.
“Mistral 7B”, 2023
Albert. Jiang et al · 2023
Cited alongside, same era.
“Large Language Models in Education: A Focus on the Complementary Relationship between Human Teachers and ChatGPT”
Jaeho Jeon and Seongyong Lee · 2023
Cited alongside, same era.
“Cybercrime and Privacy Threats of Large Language Models”
Nir Kshetri · 2023
Cited alongside, same era.
“Efficient Memory Management for Large Language Model Serving with PagedAttention”, 2023
Woosuk Kwon et al · 2023
Cited alongside, same era.
“Convergence of High Performance Computing, Big Data, and Machine Learning Applications on Containerized Infrastructures”
Peini Liu · 2023
Cited alongside, same era.
“Council Post: The Generative AI Frontier: Mastering LLM Adoption For CEOs And CTOs To Drive Business Success”
“Qwen2 Technical Report”, 2024
2024
Closest in time.
URL: https://www.mpg.de/de
“Startseite - Max-Planck-Gesellschaft”, 2024 · 2024
Closest in time.
“Llama 3 Model Card”, 2024
AI@Meta · 2024
Closest in time.
“Danny-Avila/LibreChat”, 2024
Danny Avila · 2024
Closest in time.
“Dragonfly: Multi-Resolution Zoom Supercharges Large Visual-Language Model”, 2024
Kezhen Chen et al · 2024
Closest in time.
“Software Resource Disaggregation for HPC with Serverless Computing”, 2024
Marcin Copik et al · 2024
Closest in time.
“HAWK-Digital-Environments/HAWKI”, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Manojkumar Parmar · 2023
Cited alongside, same era.
“Overcoming the Pitfalls of HPC-based Cryptojacking Detection in Presence of GPUs”
Claudius Pott, Berk Gulmezoglu and Thomas Eisenbarth · 2023
Cited alongside, same era.
“Privacy and Data Protection in ChatGPT and Other AI Chatbots: Strategies for Securing User Information”, 2023
Glorin Sebastian · 2023
Cited alongside, same era.
“OpenVenus: An Open Service Interface for HPC Environment Based on SLURM”
Meng Wan et al · 2023
Cited alongside, same era.
“CogVLM: Visual Expert for Pretrained Language Models”, 2023
Weihan Wang et al · 2023
Cited alongside, same era.
“ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats”, 2023
Xiaoxia Wu, Zhewei Yao and Yuxiong He · 2023
Cited alongside, same era.
URL: https://huggingface.co/
“Hugging Face – The AI Community Building the Future.”, 2024 · 2024
Cited alongside, same era.
HAWK Environments · 2024
Closest in time.
“Teachers’ Agency in the Era of LLM and Generative AI: Designing Pedagogical AI Agents”
Yu-Ju Lan and Nian-Shing Chen · 2024
Closest in time.
“Running Kubernetes Workloads on Rootless HPC Systems Using Slurm”
Sören Metje · 2024
Closest in time.
“Oobabooga/Text-Generation-Webui”, 2024
oobabooga · 2024
Closest in time.
“GPT-4 Technical Report”, 2024
OpenAI et al · 2024
Closest in time.
“Introducing Qwen1.5”, 2024
Qwen Team · 2024
Closest in time.
“Towards Efficient and Reliable LLM Serving: A Real-World Workload Study”, 2024
Yuxin Wang et al · 2024
Closest in time.
“FhGenie: A Custom, Confidentiality-preserving Chat AI for Corporate and Scientific Use”, 2024
Ingo Weber et al · 2024
Closest in time.
“Vision-Language Models for Vision Tasks: A Survey”, 2024
Jingyi Zhang, Jiaxing Huang, Sheng Jin and Shijian Lu · 2024
Closest in time.