REVIEW 3 cited by
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have demonstrated remarkable abilities in natural language processing. However, their deployment on resource-constrained embedded devices remains difficult due to memory and computational demands. In this paper, we present an FPGA-based accelerator designed to improve LLM inference performance on embedded FPGAs. We employ post-training quantization to reduce model size and optimize for off-chip memory bandwidth. Our design features asynchronous computation and a fully pipelined accelerator for matrix-vector multiplication. Experiments of the TinyLlama 1.1B model on a Xilinx ZCU102 platform show a 14.3-15.8x speedup and a 6.1x power efficiency improvement over running exclusively on ZCU102 processing system (PS).
Forward citations
Cited by 3 Pith papers
-
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
TeLLMe is the first edge FPGA accelerator that runs a 1.58-bit ternary LLM end-to-end, including prefill and decoding, achieving 9.51 tokens/s and 0.55 to 1.15 second prefill under 7 watts.
-
NSFlow: An End-to-End FPGA Framework with Scalable Dataflow Architecture for Neuro-Symbolic AI
NSFlow automatically generates FPGA accelerator designs for neuro-symbolic AI workloads, cutting inference latency by up to 31x versus an embedded GPU.
-
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
A custom Verilog module accelerates the token-acceptance step of speculative decoding by about 7x over GPUs, but the step is only a small fraction of total LLM inference.
Discussion (0). Continue with ORCID to comment.