Pith. sign in

REVIEW 3 cited by

Leveraging Uncertainty Estimation for Efficient LLM Routing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11021 v1 pith:O5ZUJIZN submitted 2025-02-16 cs.NI cs.CL

classification cs.NIcs.CL
keywords routingqualityresponsecostaccuracyefficiencyefficientestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deploying large language models (LLMs) in edge-cloud environments requires an efficient routing strategy to balance cost and response quality. Traditional approaches prioritize either human-preference data or accuracy metrics from benchmark datasets as routing criteria, but these methods suffer from rigidity and subjectivity. Moreover, existing routing frameworks primarily focus on accuracy and cost, neglecting response quality from a human preference perspective. In this work, we propose the Confidence-Driven LLM Router, a novel framework that leverages uncertainty estimation to optimize routing decisions. To comprehensively assess routing performance, we evaluate both system cost efficiency and response quality. In particular, we introduce the novel use of LLM-as-a-Judge to simulate human rating preferences, providing the first systematic assessment of response quality across different routing strategies. Extensive experiments on MT-Bench, GSM8K, and MMLU demonstrate that our approach outperforms state-of-the-art routing methods, achieving superior response quality while maintaining cost efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LaCy: What Small Language Models Can and Should Learn is Not Just a Question of Loss

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    LaCy augments next-token loss with spaCy parser signals to train SLMs on learn-vs-delegate decisions, producing higher FactScores in cascades than loss-only or LLM-judge baselines.

  2. Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Steering vectors flip most unjustified self-preference decisions of an LLM judge but also disturb legitimate ones, showing the bias is not captured by a single linear direction.

  3. Balancing Information Accuracy and Response Timeliness in Networked LLMs

    cs.LG 2025-08 conditional novelty 5.0 of 10

    For binary questions, combining m specialized LLMs with a Bayesian majority rule improves accuracy, and the paper derives the optimal m that trades accuracy against system delay.

Pith tools