Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have become extremely potent instruments with exceptional capacities for comprehending and producing human-like text in a wide range of applications.
Maha Elbayad et al · 1910
Earlier work this paper cites.
Iterative solution of nonlinear equations in several variables
James M Ortega and Werner C Rheinboldt. 2000 · 2000
Earlier work this paper cites.
xPilot: A Platform-Based Behavioral Synthesis System. In SRC Techcon
Deming Chen et al · 2005
Earlier work this paper cites.
Deep neural network model and FPGA accelerator co-design: Opportunities and challenges. In ICSICT
Cong Hao et al · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern et al · 2018
Earlier work this paper cites.
DNNBuilder: An automated tool for building high-performance DNN hardware accelerators for FPGAs. In ICCAD
Xiaofan Zhang et al · 2018
Earlier work this paper cites.
FPGA/DNN co-design: An efficient design methodology for IoT intelligence on the edge. In DAC
Cong Hao et al · 2019
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis et al · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu et al · 2021
Earlier work this paper cites.
MLIR: Scaling compiler infrastructure for domain specific computation. In 2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 2–14
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko. 2021 · 2021
Earlier work this paper cites.
How many layers and why? An analysis of the model depth in transformers. In IJCNLP Student Research Workshop
Antoine Simoulin et al · 2021
Earlier work this paper cites.
FPGA HLS Today: Successes, Challenges, and Opportunities
Jason Cong, Jason Lau, Gai Liu, Stephen Neuendorffer, Peichen Pan, Kees Vissers, and Zhiru Zhang. 2022 · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao et al · 2022
Earlier work this paper cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers et al · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts. In ICML
Nan Du et al · 2022
Earlier work this paper cites.
DFX: A low-latency multi-FPGA appliance for accelerating transformer-based text generation. In MICRO
Seongmin Hong et al · 2022
Earlier work this paper cites.
Confident adaptive language modeling
Tal Schuster et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei et al · 2022
Earlier work this paper cites.
AutoDistill: An end-to-end framework to explore and distill hardware-efficient language models
Xiaofan Zhang et al · 2022
Earlier work this paper cites.
Designing effective sparse expert models
Barret Zoph et al · 2022
Cited alongside, same era.
Fixing hardware security bugs with large language models
Baleegh Ahmad et al · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023 · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Charlie Chen et al · 2023
Cited alongside, same era.
Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. In ISCA
Vitcod: Vision transformer acceleration via dedicated algorithm and accelerator co-design. In HPCA
Haoran You et al · 2023
Later among the works it cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang et al · 2023
Later among the works it cites.
LLM4EDA: Emerging Progress in Large Language Models for Electronic Design Automation
Ruizhe Zhong et al · 2023
Later among the works it cites.
CHARM: C omposing H eterogeneous A ccele R ators for M atrix Multiply on Versal ACAP Architecture. In Proceedings of the 2023 ACM/SIGDA International Symposium on Field Programmable Gate Arrays . 153–164
Jinming Zhuang, Jason Lau, Hanchen Ye, Zhuoping Yang, Yubo Du, Jack Lo, Kristof Denolf, Stephen Neuendorffer, Alex Jones, Jingtong Hu, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Norm Jouppi et al · 2023
Cited alongside, same era.
AutoScaleDSE: A scalable design space exploration engine for high-level synthesis. In TRETS
Hyegang Jun et al · 2023
Cited alongside, same era.
Llm-assisted generation of hardware assertions
Rahul Kande et al · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention. In SOSP
Woosuk Kwon et al · 2023
Cited alongside, same era.
Acoustic Model Fusion For End-to-End Speech Recognition. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 1–7
Zhihong Lei, Mingbin Xu, Shiyi Han, Leo Liu, Zhen Huang, Tim Ng, Yuanyuan Zhang, Ernest Pusateri, Mirko Hannemann, Yaqiao Deng, et al · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding. In ICML
Yaniv Leviathan et al · 2023
Cited alongside, same era.
Unlocking hardware security assurance: The potential of llms
Xingyu Meng et al · 2023
Cited alongside, same era.
Using llms to facilitate formal verification of rtl
Marcelo Orenes-Vera et al · 2023
Cited alongside, same era.
Hongzheng Chen et al · 2024
Closest in time.
Break the sequential dependency of llm inference using lookahead decoding
Yichao Fu et al · 2024
Closest in time.
Albert Q Jiang et al · 2024
Closest in time.
Hydragen: High-Throughput LLM Inference with Shared Prefixes
Jordan Juravsky et al · 2024
Closest in time.
Efficiently Distilling LLMs for Edge Applications
Achintya Kundu et al · 2024
Closest in time.
Personalization of ctc-based end-to-end speech recognition using pronunciation-driven subword tokenization. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 10096–10100
Zhihong Lei, Ernest Pusateri, Shiyi Han, Leo Liu, Mingbin Xu, Tim Ng, Ruchir Travadi, Youyuan Zhang, Mirko Hannemann, Man-Hung Siu, et al · 2024
Closest in time.
SnapKV: LLM Knows What You are Looking for Before Generation
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. 2024 · 2024
Closest in time.
ChIRAAG: ChatGPT Informed Rapid and Automated Assertion Generation
Bhabesh Mali et al · 2024
Closest in time.
Towards AI-Assisted Synthesis of Verified Dafny Methods
Md Rakib Hossain Misu et al · 2024
Closest in time.
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team et al · 2024
Closest in time.
Software/Hardware Co-design for LLM and Its Application for Design Verification. In ASP-DAC
Lily Jiaxin Wan et al · 2024
Closest in time.
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
Haoyuan Wu et al · 2024
Closest in time.
HIDA: A Hierarchical Dataflow Compiler for High-Level Synthesis. In ASPLOS
Hanchen Ye et al · 2024
Closest in time.
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGA
Shulin Zeng et al · 2024
Closest in time.