Pith. sign in

REVIEW 1 cited by

Integrating NVIDIA Deep Learning Accelerator (NVDLA) with RISC-V SoC on FireSim

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.06495 v2 pith:UQYM624H submitted 2019-03-05 cs.DC cs.CV

classification cs.DCcs.CV
keywords nvdlaacceleratorfpgaperformancerunningacceleratorsdeepfiresim
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

NVDLA is an open-source deep neural network (DNN) accelerator which has received a lot of attention by the community since its introduction by Nvidia. It is a full-featured hardware IP and can serve as a good reference for conducting research and development of SoCs with integrated accelerators. However, an expensive FPGA board is required to do experiments with this IP in a real SoC. Moreover, since NVDLA is clocked at a lower frequency on an FPGA, it would be hard to do accurate performance analysis with such a setup. To overcome these limitations, we integrate NVDLA into a real RISC-V SoC on the Amazon cloud FPGA using FireSim, a cycle-exact FPGA-accelerated simulator. We then evaluate the performance of NVDLA by running YOLOv3 object-detection algorithm. Our results show that NVDLA can sustain 7.5 fps when running YOLOv3. We further analyze the performance by showing that sharing the last-level cache with NVDLA can result in up to 1.56x speedup. We then identify that sharing the memory system with the accelerator can result in unpredictable execution time for the real-time tasks running on this platform. We believe this is an important issue that must be addressed in order for on-chip DNN accelerators to be incorporated in real-time embedded systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flexible Vector Integration in Embedded RISC-V SoCs for End to End CNN Inference Acceleration

    cs.DC 2025-07 reject novelty 4.0 of 10

    Using a Hwacha vector coprocessor, the authors report up to 9x faster image preprocessing and 3x faster fallback execution for YOLOv3 on a NVDLA-based RISC-V SoC, but they mislabel Hwacha as RISC-V Vector 1.0.

Pith tools