Pith. sign in

REVIEW 5 cited by

StructuredRAG: JSON Response Formatting with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11061 v1 pith:BQVZA7H5 submitted 2024-08-07 cs.CL

classification cs.CL
keywords llmspromptingfindmodelsstrategiestasksacrossb-instruct
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability of Large Language Models (LLMs) to generate structured outputs, such as JSON, is crucial for their use in Compound AI Systems. However, evaluating and improving this capability remains challenging. In this work, we introduce StructuredRAG, a benchmark of six tasks designed to assess LLMs' proficiency in following response format instructions. We evaluate two state-of-the-art LLMs, Gemini 1.5 Pro and Llama 3 8B-instruct with 4-bit quantization using two distinct prompting strategies. We introduce these prompting strategies as f-String and Follow the Format (FF) prompting. Across 24 experiments, we find an average success rate of 82.55%. We further find a high variance in performance across tasks, models, and prompting strategies with success rates ranging from 0 to 100%. We find that Llama 3 8B-instruct often performs competitively with Gemini 1.5 Pro. We observe that task complexity significantly influences performance, with tasks involving lists or composite object outputs proving more challenging. Our findings highlight the need for further research into improving the reliability and consistency of structured output generation in LLMs. We have open-sourced our experimental code and results at github.com/weaviate/structured-rag.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Interactive Reasoning, instantiated as Hippo, lets users view and edit an LLM's chain-of-thought as a tree, and a 16-person study reports improved perceived control, sense-making, and assumption awareness.

  2. REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Asking a reasoning model several problems at once reveals large accuracy drops and exposes differences that single-question benchmarks miss.

  3. OASBuilder: Generating OpenAPI Specifications from Online API Documentation with Large Language Models

    cs.SE 2025-07 conditional novelty 5.0 of 10

    A modular pipeline of rules and LLMs converts HTML API documentation into OpenAPI specifications, evaluated on hundreds of APIs and deployed in an enterprise setting.

  4. How do Scaling Laws Apply to Knowledge Graph Engineering Tasks? The Impact of Model Size on Large Language Model Performance

    cs.AI 2025-05 conditional novelty 5.0 of 10

    On most RDF and SPARQL engineering tasks, larger open LLMs score higher, but plateau, ceiling, and occasional intra-family drops mean bigger is not always better.

  5. AI Prototyper: A Figma Plugin for Decomposition-Based GUI Prototyping with LLMs

    cs.SE 2026-07 conditional novelty 4.0 of 10

    An LLM-based Figma plugin with a human-editable feature list and RAG component retrieval produced more completed prototypes and higher expert quality ratings than manual Figma in a small Thai-language pilot.

Pith tools