Pith. sign in

REVIEW 1 cited by

Efficient Memory Management for Deep Neural Net Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.03288 v3 pith:WN6X5SGC submitted 2020-01-10 cs.LG cs.CV

classification cs.LGcs.CV
keywords memorydeepinferenceneuraldevicesefficientonlytask
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While deep neural net inference was considered a task for servers only, latest advances in technology allow the task of inference to be moved to mobile and embedded devices, desired for various reasons ranging from latency to privacy. These devices are not only limited by their compute power and battery, but also by their inferior physical memory and cache, and thus, an efficient memory manager becomes a crucial component for deep neural net inference at the edge. We explore various strategies to smartly share memory buffers among intermediate tensors in deep neural nets. Employing these can result in up to 11% smaller memory footprint than the state of the art.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FluidML: Fast and Memory Efficient Inference Optimization

    cs.LG 2024-11 reject novelty 5.0 of 10

    FluidML combines graph splitting, dynamic programming, and greedy memory allocation to optimize ML inference memory layout, but its reported improvements are inconsistent across models and its headline numbers contrad...

Pith tools