REVIEW 2 cited by
Synergy: Towards On-Body AI via Tiny AI Accelerator Collaboration on Wearables
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The advent of tiny artificial intelligence (AI) accelerators enables AI to run at the extreme edge, offering reduced latency, lower power cost, and improved privacy. When integrated into wearable devices, these accelerators open exciting opportunities, allowing various AI apps to run directly on the body. We present Synergy that provides AI apps with best-effort performance via system-driven holistic collaboration over AI accelerator-equipped wearables. To achieve this, Synergy provides device-agnostic programming interfaces to AI apps, giving the system visibility and controllability over the app's resource use. Then, Synergy maximizes the inference throughput of concurrent AI models by creating various execution plans for each app considering AI accelerator availability and intelligently selecting the best set of execution plans. Synergy further improves throughput by leveraging parallelization opportunities over multiple computation units. Our evaluations with 7 baselines and 8 models demonstrate that, on average, Synergy achieves a 23.0 times improvement in throughput, while reducing latency by 73.9% and power consumption by 15.8%, compared to the baselines.
Forward citations
Cited by 2 Pith papers
-
Test-Time Adaptation with Binary Feedback
BiTTA guides test-time model adaptation with a few binary correct/incorrect feedback labels, combining feedback-guided updates on uncertain samples with agreement-based self-adaptation on confident ones.
-
Smaller, Smarter, Closer: The Edge of Collaborative Generative AI
Edge-first collaborative inference using small language models with cloud fallback is presented as a viable design, supported by a small experiment in which load-aware scheduling halves cloud offload costs.
Discussion (0). Continue with ORCID to comment.