Pith. sign in

REVIEW 3 cited by

ElecBench: a Power Dispatch Evaluation Benchmark for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05365 v2 pith:6GBIU4GG submitted 2024-07-07 cs.AI

classification cs.AI
keywords powersectorbenchmarkelecbenchevaluationscenarioslanguagellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In response to the urgent demand for grid stability and the complex challenges posed by renewable energy integration and electricity market dynamics, the power sector increasingly seeks innovative technological solutions. In this context, large language models (LLMs) have become a key technology to improve efficiency and promote intelligent progress in the power sector with their excellent natural language processing, logical reasoning, and generalization capabilities. Despite their potential, the absence of a performance evaluation benchmark for LLM in the power sector has limited the effective application of these technologies. Addressing this gap, our study introduces "ElecBench", an evaluation benchmark of LLMs within the power sector. ElecBench aims to overcome the shortcomings of existing evaluation benchmarks by providing comprehensive coverage of sector-specific scenarios, deepening the testing of professional knowledge, and enhancing decision-making precision. The framework categorizes scenarios into general knowledge and professional business, further divided into six core performance metrics: factuality, logicality, stability, security, fairness, and expressiveness, and is subdivided into 24 sub-metrics, offering profound insights into the capabilities and limitations of LLM applications in the power sector. To ensure transparency, we have made the complete test set public, evaluating the performance of eight LLMs across various scenarios and metrics. ElecBench aspires to serve as the standard benchmark for LLM applications in the power sector, supporting continuous updates of scenarios, metrics, and models to drive technological progress and application.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Multimodal Dataset for Large Language Model Applications in the Energy Domain

    eess.SY 2026-07 accept novelty 6.5 of 10

    mAIEnergy is a harmonized, FAIR multimodal energy dataset (text, imagery, numerical series, geospatial graphs) purpose-built for LLM pre-training and retrieval-augmented generation.

  2. PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SFT plus feasibility-aware GRPO lets 3B LLMs produce electricity–computing co-schedules that are far more grid-feasible and cheaper than untrained or frontier zero-shot baselines on ECBench.

  3. VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge

    eess.SY 2026-07 conditional novelty 5.0 of 10

    Adding a power-flow solver as a tool lifts LLM accuracy on distribution OPF multiple-choice questions from 41–49% to 97–100%.

Pith tools