Pith. sign in

REVIEW 4 major objections 7 minor 23 references

SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text

T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read SAGEval claims that a role-based critiquing agent layered over a G-Eval-style scorer rectifies reference-free NLG scores and improves alignment with human annotations across all reported scoring criteria, with the largest gains on…

desk verdict SAGEval's new dataset is the real contribution; the claimed improvement from the critiquing agent is plausible but not established, and the paper needs a proper attribution control before the headline claim can be trusted. read the letter →

arxiv 2411.16077 v1 pith:3VU7RR73 submitted 2024-11-25 cs.CL cs.MA

classification cs.CLcs.MA
keywords SAGEvalreference-freeNLGevaluationLLM-as-a-judgecritiquingagentmeta-evaluationopen-endedtextgenerationsurveyhumanalignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Open-ended, reference-free text such as generated surveys, quizzes, and forms is hard to evaluate automatically because there is no gold answer to compare against, and single-pass LLM scorekeepers tend to assign inflated ratings. This paper proposes SAGEval, a two-agent setup in which a G-Eval-style Evaluator Agent scores each generated form on eight pre-defined aspects and a separate SAGE Agent critiques those scores, proposes rectified scores, suggests revisions to criterion definitions, and recommends new criteria such as creativity and content quality. The central claim is that these agent-corrected scores align better with human linguist annotations than G-Eval, CheckEval, FreeEval, MATEval, or ChatGPT-4o across the criteria that were compared; the largest gains appear on Accuracy (Spearman 0.63 vs 0.49 for G-Eval), Audience Understanding, and Audience Engagement. If the claim holds, reference-free evaluation of structured open-ended text can get closer to human judgement without additional labeled data.

What carries the argument

The load-bearing mechanism is the SAGE Agent, a second LLM prompted as a meta-evaluator. After the G-Eval-inspired Evaluator Agent assigns a 1-to-5 score with chain-of-thought reasoning and in-context exemplars, the SAGE Agent independently examines the same text and can recommend score rectifications, propose changes to the wording of aspect definitions, and suggest additional aspects. The final rectified scores are normalized using the probabilities of output tokens, following G-Eval's approach, to dampen score skew and ties.

What would settle it

Run a control where each Evaluator Agent score is decreased by a fixed amount, or by the same distribution of decreases that the SAGE Agent produces but decoupled from the specific text, and recompute the Spearman correlation with the human annotations; if this no-critique control matches or beats SAGEval, the critique content is not what is doing the work.

Watch

Extended reading notes

Core claim

The paper's discovery is that a second, role-based agent can serve as a meta-evaluator and correct the scores of a first-pass LLM evaluator in settings where no reference text exists. On a new benchmark of 96 GPT-3.5-Turbo generated surveys and forms annotated by four linguists, the SAGE Agent disagreed with the Evaluator Agent's scores in roughly 92 percent of rectifications, mostly moving scores downward from 4s and 5s to 2s and 3s. After those rectifications, SAGEval's scores showed the highest Spearman and Kendall-Tau correlations with human annotations among all compared methods, with Accuracy improving from 0.49 with G-Eval to 0.63, and Audience Engagement improving from 0.25 with G-Eval to 0.49. The SAGE Agent also proposed modified definitions for criteria such as Sentiment/Tone and suggested new criteria, most often Creativity Score and Content Quality Score, across all 96 examples. The authors interpret this as evidence that LLM evaluators can not only score but also critique and extend evaluation rubrics in the absence of ground-truth labels.

Load-bearing premise

Everything rests on treating the SAGE Agent's critique content as the cause of improved alignment, but the paper never rules out the simpler explanation that just lowering the Evaluator Agent's inflated scores, which is what 92 percent of the critiques do, is what improves the correlation.

Editorial extensions

If this is right

  • If SAGEval's claim is correct, single-agent LLM evaluation biases can be partially corrected by a second agent that critiques rather than votes.
  • Reference-free evaluation for structured open-ended outputs such as forms, surveys, and quizzes can reach human-alignment levels without labeled references.
  • The same setup can surface missing evaluation criteria, showing where initial rubrics under-specify quality.
  • Production cost stays modest relative to multi-agent debate: two LLM calls per instance are enough.
  • The released dataset and human annotations provide a reproducibility base for future reference-free evaluation work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's attribution of gains to the critique content is not fully controlled: since 92 percent of rectifications are downward, a uniform one-point lowering of the Evaluator Agent's scores might reproduce much of the correlation shift.
  • A natural extension is to apply the two-agent design to continuous reference-free text such as stories and dialogues, where rubric adaptation is even harder; the paper's own limitations section lists this as untested.
  • The 96-instance benchmark means the reported correlations carry wide uncertainty, so re-estimating them on a larger and multilingual set of generated forms would test whether the Accuracy gain of 0.14 is stable.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper introduces SAGEval, a two-agent framework for reference-free evaluation of open-ended NLG output, specifically forms/surveys generated by LLMs. An Evaluator Agent (modeled on G-Eval) assigns scores on seven predefined aspects, and a critiquing SAGE Agent provides score rectifications, suggests modifications to aspect definitions, and proposes new aspects. The authors release a dataset of 96 GPT-3.5-Turbo-generated surveys/forms annotated by four linguists, and report Spearman and Kendall correlations against these annotations, claiming that SAGEval outperforms G-Eval, CheckEval, ChatGPT-4o, FreeEval, and MATEval across all criteria. The central claim is that the critiquing mechanism yields better human alignment; the paper also presents descriptive findings about the SAGE Agent's score distributions and aspect suggestions.

Significance. If the claims were statistically robust, SAGEval would be a useful step toward reliable reference-free evaluation of structured open-ended text, an area where existing benchmarks are scarce. The release of a 96-item dataset with human annotations for forms/surveys is a valuable resource for the community. However, the current evidence is not strong enough to establish the central claim: the evaluation relies on a small sample with point estimates only, and the improvement is not convincingly attributed to the critique mechanism rather than to a general strictness shift. The contribution is therefore potentially significant but not yet demonstrated.

major comments (4)
  1. [Section 6.2, Table 2] The central claim that SAGEval "aligns better with human annotators across all scoring criteria" is not supported by inferential statistics. At N=96, the reported Spearman differences are small for several criteria (e.g., RELEV 0.48 vs. 0.47, FAIR 0.44 vs. 0.41, COH 0.41 vs. 0.32), and no confidence intervals, bootstrap intervals, permutation tests, or p-values are provided. Given that the approximate standard error of a Spearman correlation at this sample size is around 0.1, these differences are plausibly sampling noise. The authors should report interval estimates and significance tests for the difference between SAGEval and each baseline before claiming improvement "across all scoring criteria."
  2. [Sections 6.1 and 6.2, Table 1, Figure 3] The attribution of the observed improvement to the SAGE Agent's critique is not established because no control separates the effect of the critique mechanism from the effect of a general tendency toward lower scores. Table 1 and Figure 3 show that about 92% of rectifications are negative, shifting score mass from 4s/5s to 2s/3s. A plausible alternative explanation is that any instruction to be more critical, or even a fixed downward adjustment strategy, would yield higher rank correlation with human annotations on this dataset. The appropriate controls are (a) a single-pass Evaluator Agent prompted to be stricter and (b) the SAGE Agent scoring directly without first-pass scores; neither is reported. Note that a uniform score shift would not change rank correlations, so the reader's proposed control is not decisive, but the strict-prompt control would directly test whether the specific critique mechanism is responsible.
  3. [Section 4.4, Eq. (1), and Section 6.2] The method description is insufficiently detailed to reproduce the reported results. Eq. (1) is typeset incorrectly and does not specify how the probabilities p(s_i) are obtained from the LLM, nor how the SAGE Agent's rectified scores are combined with or replace the Evaluator Agent's scores to produce the final SAGEval score used in Table 2. It is unclear whether the correlations are computed on the SAGE Agent's output alone or on some aggregation. Full prompt templates, the decoding configuration, and the exact score-rectification procedure should be provided.
  4. [Section 6.3, Figure 4] The third contribution, that the SAGE Agent can propose new scoring aspects, is presented without validation. The paper reports that the SAGE Agent suggests aspects such as "Creativity Score" and "Content Quality Score" across many data points, but it does not evaluate whether these suggested aspects improve correlation with human judgments or add useful information beyond the predefined set. Without such analysis, this claim remains descriptive and does not support the stated contribution of demonstrating capabilities to propose useful new aspects.
minor comments (7)
  1. [Abstract] The abstract contains typos: "more more complex" should be "more complex," and "doesn't exist or isn't amply available" is informal.
  2. [Section 6.3] The sentence beginning "SAGE Agent is prompted to also ensure if the pre-defined scoring criteria is" is incomplete; the paragraph appears to be cut off.
  3. [Table 1] The column headers "Neg Pos Total Definition" are confusing. The "Definition" column header does not match its content (counts of instances), and "Neg" and "Pos" should be spelled out or defined in the caption.
  4. [Section 6.2] The text refers to "LLMEval (which is based on G-EVal framework)" but Table 2 labels the row "G-Eval"; use consistent terminology throughout.
  5. [Section 5] The sentence "the annotated by 4 experienced linguists" is missing a verb; it should read "annotated by 4 experienced linguists."
  6. [Section 2] There is a typo in "multi-agent farmeworks" which should be "frameworks."
  7. [Figure 3 caption] The caption "Scores distribution by SAGE Agent compared scores assigned by Evaluator Agent" is ungrammatical; consider "Score distributions assigned by the SAGE Agent and the Evaluator Agent."

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: SAGEval is an empirical framework evaluated against external human annotations; no fitted parameter is renamed as a prediction and the one overlapping-author citation is incidental.

full rationale

The paper's central claim is that a critiquing SAGE Agent rectifies scores produced by a G-Eval-style Evaluator Agent and that the rectified scores correlate better with human judgments on a newly collected dataset of 96 reference-free surveys/forms. The derivation chain is empirical rather than definitional: the Evaluator Agent scores each instance on eight predefined aspects, the SAGE Agent proposes score corrections and optionally suggests new aspects, and the resulting scores are compared with human annotations from four linguists. No score is fitted to the human labels, no parameter is estimated from the target correlations, and the human annotations are external to the framework's prompts. The aspect definitions are explicitly adapted from prior work (GPT-Score and X-eval), and the G-Eval formulation is cited as an external baseline rather than invoked to forbid alternatives. The only self-citation involving a current author, Wadhwa et al. (2024), appears in passing in the introduction as an example of retrieval-augmented generation and is not load-bearing for any of the paper's evaluation claims. The absence of confidence intervals and attribution controls (e.g., a strict-scoring single-pass evaluator) is a statistical and experimental-design limitation, not a circularity: the framework's output is not equivalent to its input by construction, and the comparison with human annotations gives the claim independent empirical content. No step in the paper reduces an equation to itself, no fitted value is renamed as a prediction, and no load-bearing premise rests on a self-citation chain. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The ledger is mostly empty of fitted parameters and invented entities; the main dependencies are domain assumptions about LLM judgment, the reliability of the small human-labeled dataset, and the transfer of G-Eval's normalization. The 8 predefined aspects are adopted from prior work, not fitted here.

assumptions (5)
  • domain assumption LLM evaluators (GPT-4) can assign meaningful 1-5 scores on the predefined 8 aspects for open-ended forms.
    The entire evaluation framework relies on this capability, assumed throughout Sections 4 and 6.
  • domain assumption The SAGE Agent's score rectifications and definition suggestions reflect valid reasoning rather than artifacts of prompt phrasing or output distribution.
    Without this, the claimed improvement could be an artifact; the paper does not isolate this assumption.
  • domain assumption Human annotations from 4 linguists are a reliable gold standard for the 8 aspects.
    The paper treats human scores as ground truth but reports no inter-annotator agreement.
  • standard math The probability normalization in Equation (1) from G-Eval is a valid way to aggregate token probabilities into scores.
    Taken from G-Eval (Liu et al., 2023b), used as is in Section 4.4.
  • domain assumption The 96 GPT-3.5-generated surveys are representative of open-ended reference-free structured text.
    The paper generalizes beyond this specific dataset without evidence of representativeness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text." pith.science (2026). https://pith.science/paper/3VU7RR73

@misc{pith2026241116077,
  author       = {Pith},
  title        = {Pith review of: SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3VU7RR73}},
  note         = {Machine review of arXiv:2411.16077}
}
read the original abstract

Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity and time savings. But as these integrations become more more complex, it is paramount to ensure that the quality of output from the LLM-integrated applications are relevant and appropriate for use. Identifying the need to develop robust evaluation approaches for natural language generation, wherein references/ground labels doesn't exist or isn't amply available, this paper introduces a novel framework called "SAGEval" which utilizes a critiquing Agent to provide feedback on scores generated by LLM evaluators. We show that the critiquing Agent is able to rectify scores from LLM evaluators, in absence of references/ground-truth labels, thereby reducing the need for labeled data even for complex NLG evaluation scenarios, like the generation of JSON-structured forms/surveys with responses in different styles like multiple choice, likert ratings, single choice questions, etc.

Figures

Figures reproduced from arXiv: 2411.16077 by the authors.

Figure 1
Figure 1. SAGEval framework. SAGEval engages with a "wiser" role-based agent to validate scores assigned by the first LLM Evaluator for reference-free texts. have sub-optimal results only using a prompted call to an LLM. This popularity has enabled the exploration of agentic frameworks in LLM-based automatic evaluation of natural language text (Chan et al., 2023; Li et al., 2024), but these approaches still do not solve/exami… view at source ↗
Figure 2
Figure 2. Open-ended human-drafted and NLG texts like lists, surveys, forms, contains sub-items or entities that are associated with a central theme such as "List of things to pack while traveling", or "Survey on assessing the quality of healthcare services", but these items (bullets in a list, questions in a survey) differ from each other, and it is important to make sure that the variance in open-ended text is coherent and … view at source ↗
Figure 3
Figure 3. Scores distribution by SAGE Agent compared scores assigned by Evaluator Agent. We find that Evaluator Agent is inclined towards assigning higher ratings (4s and 5s) across all criteria, whereas SAGE Agent is more critical and pushes the score distribution towards 3s and a couple of 2s. Scoring Criteria ACC SEMD COH RELEV AUND AENG FAIR ρ τ ρ τ ρ τ ρ τ ρ τ ρ τ ρ τ G-Eval 0.49 0.44 0.62 0.49 0.32 0.43 0.47 0.43 0.33 0… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Term-topic frequency distributions of sug [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Distribution of annotation scores (between 1-5) assigned to each Scoring Criteria: , by 4 highly experienced [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

23 extracted references · 18 canonical work pages

  1. [1]

    Read the generated output form/survey/quiz text carefully and identify the main theme across all sections and questions, and option choices (for the case of multichoice and single choice questions)

  2. [2]

    arXiv preprint arXiv:2403.00839

    Toolnet: Connecting large language models with massive tools via tool graph. arXiv preprint arXiv:2403.00839. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. Gpte- val: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634. Tinh Son Luong, Thanh-Thien Le, Linh Ngo Van, and Thien Huu Ng...

  3. [3]

    arXiv preprint arXiv:2404.13076

    Llm evaluators recognize and favor their own generations. arXiv preprint arXiv:2404.13076. Kishore Papineni, Salim Roukos, Todd Ward, and Wei- Jing Zhu. 2002. Bleu: a method for automatic evalu- ation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computa- tional Linguistics, pages 311–318. Chengwei Qin, Aston Zha...

  4. [5]

    Ad- vances in Neural Information Processing Systems , 36

    How2comm: Communication-efficient and collaboration-pragmatic multi-agent perception. Ad- vances in Neural Information Processing Systems , 36. Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023. Mm-vet: Evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.0...

  5. [7]

    Check if the general theme of the content in form/survey/quiz is aligned to the theme of the prompt (user ask), and if it presents them in a clear and logical order

  6. [8]

    Assign a score for Accuracy on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Semantic Diversity: This criteria looks at the generated output text and then tries to judge whether the questions across all sections (if present) and the form are diverse, meaning they are semantically different and there are no...

  7. [9]

    Read the generated output form/survey/quiz text carefully and ensure that there are no duplicates

  8. [10]

    Also check if the content in form/survey/quiz is semantically rich and aligns to the theme of the prompt (user ask), while being diverse/different from each other

Show all 23 references
  1. [11]

    Assign a score for Semantic Diversity on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Cohension: This criteria looks at the gener- ated output text and then tries to judge whether the questions across all sections (if present)...

  2. [12]

    Read the generated output form/survey/quiz text carefully and ensure that there are no typos or grammatical errors. 2. Also check if the content in form/survey/quiz is fluent in english and coherent to understand

  3. [13]

    Assign a score for Cohesion on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Relevancy: Evaluation on fourth Criteria: This criteria looks at the generated output text and then tries to judge whether the questions across all se...

  4. [14]

    user ask

    Read the generated output form/survey/quiz text carefully and ensure that all questions, section titles and options are relevant and important to the "user ask"

  5. [15]

    Assign a score for Relevancy on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Audience Understandability: This criteria looks at the generated output text and then tries to judge whether the questions across all sections (if pr...

  6. [17]

    Audience Understandability

    After reading through the contents of the form/survey/quiz generated, please assign a "Audience Understandability" score on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Audience Engagement score: This criteria looks at the gen...

  7. [18]

    Assume that you (GPT4 model) are the responder of the form/survey/quiz generated, and now read the generated output form/survey/quiz text carefully

  8. [19]

    Audience Engagement

    After reading through the contents of the form/survey/quiz generated, please assign a "Audience Engagement" score on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Fairness score: This criteria looks at the generated output text...

  9. [20]

    Read the generated output form/survey/quiz text carefully and ensure that all questions, section titles, title of the form, description of the form are generated in a language that is fair, without any bias, or harmful content, that may cause discomfort to the responders

  10. [21]

    Also check if the conetnt in form/survey/quiz should be flagged on any Responsible AI stan- dards

  11. [22]

    Assign a score for Fairness on a scale of 1 to 5, where 1 is the lowest and 5 is the highest based on the Evaluation Criteria. Sentiment/Tone type: This criteria looks at the generated output text and then tries to identify the sentiment of the content by analyzing the questio...

  12. [23]

    Read the generated output form/survey/quiz text carefully and identify from the language of all questions, section titles, title of the form, description of the form the sentiment it conveys

  13. [24]

    Unlike the previous evaluation criteria which were assign a score for Fairness on a scale of 1 to 5, here please output tone/sentiment of the generated content (questions). """

  14. [2023]

    arXiv preprint arXiv:2310.15123

    Branch-solve-merge improves large language model evaluation and generation. arXiv preprint arXiv:2310.15123. Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini:...

  15. [2024]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al

    Welcome to copilot in microsoft forms. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Kabir Ahuj...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.