Pith. sign in

REVIEW 12 cited by

HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.16191 v1 pith:NSBEXWHA submitted 2024-09-24 cs.CL

classification cs.CL
keywords textlonggenerationllmsevaluationcapabilitieshellobenchhelloeval
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks (e.g., long-context understanding), and many benchmarks have been proposed. However, we observe that long text generation capabilities are not well investigated. Therefore, we introduce the Hierarchical Long Text Generation Benchmark (HelloBench), a comprehensive, in-the-wild, and open-ended benchmark to evaluate LLMs' performance in generating long text. Based on Bloom's Taxonomy, HelloBench categorizes long text generation tasks into five subtasks: open-ended QA, summarization, chat, text completion, and heuristic text generation. Besides, we propose Hierarchical Long Text Evaluation (HelloEval), a human-aligned evaluation method that significantly reduces the time and effort required for human evaluation while maintaining a high correlation with human evaluation. We have conducted extensive experiments across around 30 mainstream LLMs and observed that the current LLMs lack long text generation capabilities. Specifically, first, regardless of whether the instructions include explicit or implicit length constraints, we observe that most LLMs cannot generate text that is longer than 4000 words. Second, we observe that while some LLMs can generate longer text, many issues exist (e.g., severe repetition and quality degradation). Third, to demonstrate the effectiveness of HelloEval, we compare HelloEval with traditional metrics (e.g., ROUGE, BLEU, etc.) and LLM-as-a-Judge methods, which show that HelloEval has the highest correlation with human evaluation. We release our code in https://github.com/Quehry/HelloBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A task-adaptive rubric-selection method using Bayesian measurability and IRT-based greedy assembly that compresses rubric banks while improving agreement and preserving rank fidelity.

  2. RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review

    cs.CL 2026-06 conditional novelty 6.0 of 10

    A rubric-first LLM pipeline that splits peer review into rubric generation, rubric-conditioned review writing, and final scoring outperforms existing AI reviewers on alignment with human judgments in a 200-paper test.

  3. MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems

    cs.LG 2025-10 unverdicted novelty 6.0 of 10

    MemoryBench shows that state-of-the-art LLM memory systems do not reliably learn from simulated user feedback and are often outperformed by naive RAG.

  4. Reverse-Engineered Reasoning for Open-Ended Generation

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Given a high-quality output, the authors search for a thinking trace that minimizes that output's perplexity, then fine-tune Qwen3-8B on 20,000 such traces, reporting writing performance near GPT-4o and Claude 3.5.

  5. From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    ProxyReward trains long-form generation models by rewarding how well an AI judge can answer generated yes/no questions about the response, improving open-source models on ProxyQA.

  6. SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 7B writing model trained on plan-write-refine thinking data with multi-stage preference optimization matches or beats several larger models on long-form generation benchmarks.

  7. EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 728-prompt multi-genre Chinese essay benchmark whose hierarchical, genre-specific LLM scoring aligns with human rankings better than a general-purpose writing baseline, especially with DeepSeek-R1 as judge.

  8. LIFEBench: Evaluating Length Instruction Following in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LIFEBench's evaluation of 26 LLMs shows most follow short length instructions but degrade sharply beyond a few hundred words, and none reliably hit vendor-claimed maximum output lengths.

  9. FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

    cs.CL 2026-07 conditional novelty 5.5 of 10

    A multi-LLM consensus pipeline turns 14,450 auto-generated candidate rubrics into 2,600 distinguishable gold rubrics that rank 10 financial deep-research systems from 58.58% to 22.23% pass rate.

  10. SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    SGSimEval is a multifaceted benchmark showing that automatic survey generation systems match humans on outline quality but lag on content and references.

  11. MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    MiniLongBench, a 237-sample compression of LongBench, is claimed to reproduce model rankings with a 0.97 Spearman correlation at 4.5% of the evaluation cost.

  12. DynamicBench: Evaluating Real-Time Report Generation in Large Language Models

    cs.LG 2025-06 reject novelty 4.0 of 10

    DynamicBench is a proposed benchmark for real-time report generation, and the authors claim their retrieval-augmented system outperforms GPT-4o, but the evidence is inconsistent and the comparison is not apples-to-apples.

Pith tools