Pith. sign in

REVIEW 3 cited by

The Fellowship of the LLMs: Multi-Model Workflows for Synthetic Preference Optimization Dataset Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08688 v7 pith:65RY34BL submitted 2024-08-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords generationresponseevaluationllmstextitconfigurationsdatasetdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

This paper presents a novel methodology for generating synthetic Preference Optimization (PO) datasets using multi-model workflows. We evaluate the effectiveness and potential of these workflows in automating and enhancing the dataset generation process. PO dataset generation requires two modules: (1) $\textit{response evaluation}$, and (2) $\textit{response generation}$. In the $\textit{response evaluation}$ module, the responses from Large Language Models (LLMs) are evaluated and ranked - a task typically carried out by human annotators that we automate using LLMs. We assess the response evaluation module in a 2 step process. In step 1, we assess LLMs as evaluators using three distinct prompting strategies. In step 2, we apply the winning prompting strategy to compare the performance of LLM-as-a-Judge, LLMs-as-a-Jury, and LLM Debate. Our evaluation shows that GPT-4o-as-a-Judge is more consistent across all datasets. For the $\textit{response generation}$ module, we use the identified LLM evaluator configuration and compare different configurations of the LLM Feedback Loop. We use the win rate to determine the best multi-model configuration for generation. Experimenting with various configurations, we find that the LLM Feedback Loop, with Llama as the generator and Gemma as the reviewer, achieves a notable 71.8% and 73.8% win rate over single-model Llama and Gemma, respectively. After identifying the best configurations for both modules, we generate our PO datasets using the above pipeline.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A reward-oriented influence scoring method that uses pairwise preference loss to select 5% of instruction-tuning data, outperforming similarity-based selection baselines on SHP, SE, and HH-RLHF.

  2. TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

    cs.IR 2025-08 conditional novelty 5.0 of 10

    TalkPlayData 2 is a new synthetic conversational music recommendation dataset created by four cooperating multimodal LLM agents, featuring user profiles, conversation goals, chain-of-thought annotations, and cold-star...

  3. MAG-V: A Multi-Agent Framework for Synthetic Data Generation and Verification

    cs.CL 2024-11 conditional novelty 5.0 of 10

    MAG-V generates synthetic queries with LLM agents and verifies assistant tool-call trajectories using reverse-engineered questions and simple similarity features, matching GPT-4 judge accuracy.

Pith tools