Pith. sign in

REVIEW 3 cited by

Democratizing LLMs: An Exploration of Cost-Performance Trade-offs in Self-Refined Open-Source Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.07611 v2 pith:KZKQF4OZ submitted 2023-10-11 cs.CL cs.AIcs.PF

classification cs.CLcs.AIcs.PF
keywords performancemodelsllmsaccesscostcostsdemocratizinghigh-performing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The dominance of proprietary LLMs has led to restricted access and raised information privacy concerns. High-performing open-source alternatives are crucial for information-sensitive and high-volume applications but often lag behind in performance. To address this gap, we propose (1) A untargeted variant of iterative self-critique and self-refinement devoid of external influence. (2) A novel ranking metric - Performance, Refinement, and Inference Cost Score (PeRFICS) - to find the optimal model for a given task considering refined performance and cost. Our experiments show that SoTA open source models of varying sizes from 7B - 65B, on average, improve 8.2% from their baseline performance. Strikingly, even models with extremely small memory footprints, such as Vicuna-7B, show a 11.74% improvement overall and up to a 25.39% improvement in high-creativity, open ended tasks on the Vicuna benchmark. Vicuna-13B takes it a step further and outperforms ChatGPT post-refinement. This work has profound implications for resource-constrained and information-sensitive environments seeking to leverage LLMs without incurring prohibitive costs, compromising on performance and privacy. The domain-agnostic self-refinement process coupled with our novel ranking metric facilitates informed decision-making in model selection, thereby reducing costs and democratizing access to high-performing language models, as evidenced by case studies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Data-Centric Approach for Safe and Secure Large Language Models against Threatening and Toxic Content

    cs.CR 2025-04 reject novelty 4.0 of 10

    A BART-based post-generation corrector lowers toxicity and jailbreaking scores, but the reported gains are partly in-sample because thresholds are optimized on the evaluation data.

  2. A Functional Software Reference Architecture for LLM-Integrated Systems

    cs.SE 2025-01 conditional novelty 4.0 of 10

    A preliminary four-layer functional reference architecture for LLM-integrated systems, illustrated on three open-source projects.

  3. Resource-Efficient Language Models: Quantization for Fast and Accessible Inference

    cs.AI 2025-05 unverdicted

    A survey of post-training quantization techniques for large language models, covering schemes, granularities, and popular methods, with no new experimental results.

Pith tools