Pith. sign in

REVIEW 7 cited by

Platypus: Quick, Cheap, and Powerful Refinement of LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.07317 v2 pith:UIWDAWOY submitted 2023-08-14 cs.CL

classification cs.CL
keywords llmsplatypusdataopenachievesdatasetfamilyfine-tuned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We present $\textbf{Platypus}$, a family of fine-tuned and merged Large Language Models (LLMs) that achieves the strongest performance and currently stands at first place in HuggingFace's Open LLM Leaderboard as of the release date of this work. In this work we describe (1) our curated dataset $\textbf{Open-Platypus}$, that is a subset of other open datasets and which $\textit{we release to the public}$ (2) our process of fine-tuning and merging LoRA modules in order to conserve the strong prior of pretrained LLMs, while bringing specific domain knowledge to the surface (3) our efforts in checking for test data leaks and contamination in the training data, which can inform future research. Specifically, the Platypus family achieves strong performance in quantitative LLM metrics across model sizes, topping the global Open LLM leaderboard while using just a fraction of the fine-tuning data and overall compute that are required for other state-of-the-art fine-tuned LLMs. In particular, a 13B Platypus model can be trained on $\textit{a single}$ A100 GPU using 25k questions in 5 hours. This is a testament of the quality of our Open-Platypus dataset, and opens opportunities for more improvements in the field. Project page: https://platypus-llm.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

    cs.CR 2025-10 conditional novelty 6.0 of 10

    A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.

  2. Sensitivity-LoRA: Low-Load Sensitivity-Based Fine-Tuning for Large Language Models

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Sensitivity-LoRA allocates LoRA ranks across layers using Hessian-based sensitivity metrics, improving average GLUE score by 0.74 over AdaLoRA on RoBERTa-base.

  3. MISCON: A Mission-Driven Conversational Consultant for Pre-Venture Entrepreneurs in Food Deserts

    cs.AI 2025-01 conditional novelty 5.0 of 10

    MISCON is a conversational AI system that guides aspiring food business owners in food deserts through market, finance, and permit decisions using a knowledge graph and LLMs.

  4. Improving Mathematical Reasoning Capabilities of Small Language Models via Feedback-Driven Distillation

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A feedback-driven distillation pipeline iteratively generates harder variants of problems small models solve and similar problems for ones they miss, improving their math reasoning scores.

  5. Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report

    cs.CR 2025-08 conditional novelty 4.0 of 10

    Foundation-Sec-8B-Instruct, an instruction-tuned 8B cybersecurity LLM, is released and claimed to beat Llama 3.1-8B-Instruct on CTIBench-RCM and CTIBench-MCQA while remaining competitive on general instruction-following.

  6. Parameter Efficient Fine-Tuning for Deep Learning-Based Full-Waveform Inversion

    cs.CE 2024-12 conditional novelty 4.0 of 10

    LoRA fine-tuning of a pretrained InversionNet model on OpenFWI matches full fine-tuning in-distribution and improves out-of-distribution generalization for seismic full-waveform inversion.

  7. Understanding Hidden Computations in Chain-of-Thought Reasoning

    cs.CL 2024-12 reject novelty 4.0 of 10

    Hidden reasoning tokens in a transformer trained on filler chain-of-thought can reportedly be recovered as rank-2 predictions during decoding.

Pith tools