{"id":"e5b4944d-9745-4b87-b9c5-5ba7931873b6","arxiv_id":"2504.20010","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A retrieval-augmented LLM pipeline, the Problem Scoping Agent, generates AI4SG project proposals that blind human reviewers score as comparable to expert-written proposals.","lead":"The authors propose a Problem Scoping Agent that uses an LLM, web searches, and academic paper lookups to generate project proposals for AI-for-social-good organizations from just an organization's name. A blind review found the generated proposals scored similarly to (and sometimes better than) expert-written proposals that had been reformatted by an AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'comparable to experts' claim rests on a baseline transformed by GPT-4o; no evidence shows this rewrite preserves the quality or content of the original expert summaries.","rationale":"The reader's weakest assumption identifies the main load-bearing point. The experiment is set up to compare PSA output to 'expert' proposals, but the expert arm is GPT-4o-rewritten DSSG summaries (Section 4.1). The rewriting prompt is the same as the solution-generation prompt in Appendix A, which asks for a title, problem statement, and proposed solution in a specific format; this is not a neutral formatting operation. Because all statistical comparisons in Table 2 are against these rewritten originals, the abstract's 'comparable to experts' is not directly supported. For GPT-PSA, the baseline is also GPT-4o-processed, potentially creating a style and format advantage. No evidence is provided that the rewrite preserves quality. The paper is otherwise transparent: LLM-based evaluation is reported as unsatisfactory, the sample size is small, and the authors describe limitations. The concrete test of adding original unmodified summaries as an evaluation arm would settle the issue. If the originals and rewritten versions score the same, the concern evaporates; if not, the headline claim needs to be weakened to 'comparable to GPT-4o-rewritten summaries' or the rewrites need to be validated. This matches the reader's conditional verdict, and no additional load-bearing concern outweighs it.","tokens_in":13319,"tokens_out":4422,"duration_ms":45263,"concrete_test":"Re-run the blind human evaluation with the original, unmodified DSSG project summaries as an additional baseline arm alongside the GPT-4o-rewritten versions and the PSA proposals, using the same rubric, the same permutation protocol, and a pre-registered paired comparison of PSA outputs against the original expert text. If the mean difference between PSA and original expert summaries is significantly worse than the difference reported in Table 2, or if the rewritten originals score significantly lower than the originals, then the 'comparable to experts' claim is not supported; if the two baselines are statistically indistinguishable, the concern is resolved. Report the number of reviewers and inter-rater reliability so the statistical comparison is interpretable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's only human comparison is against the 'Original' row of Table 2, which is not the untouched DSSG expert text. Section 4.1 states that GPT-4o was used to rewrite the original proposals to ensure consistent formatting, and that the rewriting prompt is the same as the final proposal generation prompt in Appendix A. That prompt instructs the model to produce a title, problem statement, and proposed solution in a specific concise format. Running expert summaries through this prompt is not a formatting-only operation: it can select, rephrase, or drop content, and in particular can change concreteness, thoroughness, and scope, which are exactly the dimensions evaluated. The abstract's central claim ('comparable to those written by experts') is therefore supported only relative to GPT-4o-processed versions of expert text. If the rewrite degrades the originals, the comparison is biased in favor of the PSA. Moreover, for GPT-PSA the baseline and the agent share the same underlying model, so formatting and style similarity may further inflate scores. The paper provides no validation (e.g., human ratings of original vs rewritten text, content overlap measures) that the rewrite is lossless. This is not a minor issue: every quantitative comparison in Section 4.2 uses the rewritten row as reference, so the headline result is conditional on an unverified equivalence assumption. The paper is honest about its exploratory nature and about the failure of LLM evaluators, but the 'comparable to experts' wording overstates what is measured. The small N=21 and unspecified reviewer details compound the problem, but the rewrite issue is the load-bearing concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a Problem Scoping Agent (PSA) that combines retrieval-augmented LLM steps—background, challenge, method retrieval, and solution generation—to produce AI4SG project proposals. The authors evaluate PSA against 21 DSSG project summaries using blind human review and LLM-based evaluators, and report that DS-PSA and GPT-PSA are statistically indistinguishable from the GPT-4o-rewritten originals, that G-PSA outperforms them, and that PSA increases problem diversity relative to base LLMs. The paper also documents the failure of LLM evaluators and makes its prompts publicly available.","tokens_in":13556,"tokens_out":4847,"duration_ms":44156,"significance":"If the core claim holds, the paper makes a useful contribution to AI4SG: it formalizes a scoping pipeline and demonstrates a path toward automating a labor-intensive task. The paper is transparent about the failure of LLM evaluators, uses real-world DSSG projects, and reports diversity gains. However, the headline claim of 'comparable to experts' is presently conditional on an unvalidated rewriting step, and the human evaluation reporting is incomplete, so the empirical support is weaker than the abstract suggests.","major_comments":[{"comment":"The human baseline used for all comparisons is not the original DSSG expert text but a GPT-4o-rewritten version, created with the same prompt used for final proposal generation (Appendix A). That prompt imposes content constraints—title, problem statement, proposed solution, avoidance of trivial or outreach solutions—that can alter exactly the dimensions evaluated (appropriateness, thoroughness, feasibility, expected effectiveness). The manuscript provides no evidence that the rewriting is lossless, such as human ratings of original vs. rewritten text or a content-overlap measure. Every p-value in Table 2 is against this transformed baseline, so the abstract's claim of 'comparable to those written by experts' is supported only relative to GPT-4o-processed versions. This is a load-bearing issue and needs to be addressed, either by using the original summaries directly or by validating the rewrite and reporting original-vs-rewrite scores.","section":"Section 4.1, Table 2"},{"comment":"The human evaluation is under-specified: the number of evaluators, their qualifications, the rating procedure, and inter-rater reliability (e.g., Krippendorff's alpha or ICC) are not reported. With 21 proposals and Likert-scale ratings, this information is necessary to interpret the paired t-tests and Hotelling's T2 tests; if a single rater scored all proposals, the independence assumption underlying the tests is violated, and the reported p-values are not trustworthy.","section":"Section 4.1, 4.2"},{"comment":"The conclusion that DS-PSA and GPT-PSA produce proposals 'comparable' to expert-written ones is based on failure to reject the null hypothesis (p ≈ 0.66 and 0.70 for the average). With n = 21, this is not evidence of equivalence; the tests may simply lack power. The paper should either report confidence intervals on the mean differences or use an equivalence testing approach with a pre-specified boundary, and should discuss the minimum detectable effect. This directly affects the main claim.","section":"Section 4.2, Table 2"},{"comment":"The paper honestly reports that AI-based evaluation was unsatisfactory (low variance, low correlation with human evaluation). However, the abstract states the framework was validated 'through a blind review and AI evaluations.' Since the AI evaluations are shown to be unreliable and are relegated to the appendix, the abstract should be revised to attribute the main evidence to blind human review, or the AI evaluation should be omitted from the central claim.","section":"Section 4.2, Table 1"}],"minor_comments":[{"comment":"The notation 'µ±σ2' is ambiguous; using '±' with variance is nonstandard and may mislead readers. Consider reporting standard deviations or clarifying the header.","section":"Tables 2 and 3"},{"comment":"The footnote says all p-values are from Hotelling's T2 test, but the text says per-metric p-values are from paired t-tests. Align the footnote with the text to avoid confusion.","section":"Table 2 footnote"},{"comment":"The running example is the Memphis Fire Department, but the figure lists IBM, Conservation International, Pew Trusts, and UCSB as the organizations being searched; this is confusing and should be clarified.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is an interesting proof-of-concept, but the evaluation gap is substantial. The authors should be given the opportunity to add the original baseline comparison or validate the rewriting step; this is addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe short version: this is a genuine proof-of-concept for a new task—AI4SG problem scoping—and the PSA pipeline is a reasonable integration of retrieval and verbalized confidence. But the headline claim that PSA generates proposals 'comparable to those written by experts' is not actually measured. The comparison in Table 2 is against GPT-4o-rewritten versions of the DSSG expert summaries, not the original expert text, and the paper provides no evidence that the rewrite is lossless. Since the rewriting prompt is identical to the final generation prompt, the baseline has already been reshaped toward the exact format the PSA outputs. That is a load-bearing flaw, not a nitpick. The stress-test note is right.\n\nWhat the paper does well: it defines a task with real practical value, uses real DSSG projects as the anchor, and is honest about the failure of LLM-based evaluators (correlations near zero, low variance in AI scores). The qualitative findings—models losing objectives, lack of diversity in generated problems—are useful. The modular design (background, challenge, method retrieval) is sensible, and the verbalized-confidence weighting is a reasonable touch.\n\nSoft spots beyond the baseline rewrite: the human evaluation details are missing (number of reviewers, expertise, inter-rater reliability), N=21 is small, and no code or data are released. The tables report mean ± variance rather than standard deviation, which is odd. None of these are fatal alone, but together with the rewrite issue the central empirical claim is under-supported. The paper's own contribution (3) is more careful—'not too far from human experts in many cases'—and that is closer to what the data show.\n\nWho should read it: anyone working on LLM agents for public-sector or social-good applications, and people studying AI-assisted idea generation. It is a useful systems paper, not a fundamental advance.\n\nRecommendation: I would send it to peer review, but flag the baseline as the main revision. The authors need to compare against the untouched expert summaries, or validate that the GPT-4o rewrite preserves content and quality. They should also report reviewer numbers and reliability. If those changes are made, the 'comparable to experts' wording can be defended. As is, the abstract overstates the result.","headline":"Useful proof-of-concept for LLM-based AI4SG scoping, but the 'comparable to experts' claim rests on a GPT-4o-rewritten baseline and needs direct validation against the original expert text.","tokens_in":14156,"tokens_out":3335,"would_cite":false,"duration_ms":30667,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A retrieval-augmented LLM agent can write AI-for-social-good proposals that match expert scoping.","keywords":["AI for Social Good","problem scoping","large language models","retrieval-augmented generation","LLM agents","project proposal generation","human evaluation","LLM evaluation"],"falsifier":"Have a new panel of blinded reviewers score the original, unrewritten expert summaries alongside PSA proposals with the same four-metric rubric; if the agent no longer matches the baseline, the claimed parity is an artifact of the rewriting step. A smaller test is to score original versus rewritten expert summaries directly; material score differences mean the baseline itself was moved by the rewrite.","tokens_in":13038,"feed_emoji":"🤖","tokens_out":9207,"duration_ms":83770,"temperature":0.7,"pith_summary":"This paper contends that a critical bottleneck in AI for Social Good—turning an organization's pain point into a concrete, technically grounded project proposal—can be substantially automated with a retrieval-augmented LLM agent. The authors build the Problem Scoping Agent (PSA), which searches for an organization's background, finds current challenges relevant to it, pulls academic papers on candidate AI methods, and only then writes the proposal. Blind human review using four rubric metrics shows the resulting proposals scoring close to expert-written project summaries: DS-PSA and GPT-PSA are statistically indistinguishable from the expert baseline, and the Gemini-based G-PSA is rated higher on average. The paper also finds that LLM-as-judge evaluation is too low-variance and too weakly correlated with humans to replace human review. If the claim holds, organizations without in-house technical teams could cheaply generate many scoping options before committing resources.","feed_headline":"AI agent matches human experts at scoping social-good projects","feed_subtitle":"A retrieval-grounded LLM pipeline produced proposals blind reviewers scored comparably to expert-written ones.","key_machinery":"The load-bearing mechanism is the multi-stage retrieval and grounding loop of the PSA. Starting from an organization name, the agent annotates and summarizes retrieved web pages into a background statement, uses that to generate challenge-specific search queries, reranks candidate challenges by verbalized confidence and tractability, then queries a scholarly literature API for methods tied to the selected challenge. A divergent-convergent step—search broadly, prune to the five highest-confidence papers—breaks the LLM's tendency to fixate on generic problems, while confidence-weighted sampling of challenges promotes diversity. The final proposal prompt combines the background, the chosen challenge, and the pruned method set, forcing the model to justify each AI technique against the problem's constraints.","core_discovery":"The central claim is that a pipeline interleaving web and scholarly retrieval with LLM summarization can turn an organization name into a full AI-for-social-good project proposal that human experts rate about as well as a proposal produced by a rigorous expert scoping process. The PSA retrieves and annotates background pages on the organization, generates and searches five candidate challenges, scores each by verbalized confidence and tractability, samples one challenge with softmax-weighted probability, retrieves up to ten scholarly papers on candidate methods, prunes to the five most applicable, and generates a solution grounded in the selected challenge and methods. On 21 expert-written fellowship project summaries, averaged human Likert scores for DS-PSA and GPT-PSA were not significantly different from the baseline, while G-PSA exceeded the baseline by 0.5476 points on average (p = 0.0058). The paper is explicit that the comparison baseline is the expert summaries after a GPT-4o formatting rewrite, and that a vanilla DeepSeek-V3 without retrieval scaffolds scores 0.619 below the baseline.","pith_inferences":["Placing PSA inside a human-in-the-loop workflow could yield the practical gain the paper only gestures at: the agent drafts many cheap proposals and scarce experts spend their time filtering and refining rather than writing from scratch.","The rewrite assumption invites a direct test: scoring original, unrewritten expert summaries against PSA proposals would reveal whether the claimed parity is an artifact of GPT-4o's formatting pass.","The same retrieve-annotate-retrieve-generate loop could transfer to other under-resourced planning tasks, such as grant writing or policy briefs, whenever an organization's public record and relevant literature are available to ground the LLM."],"forward_implications":["For a weaker base model such as DeepSeek-V3, adding PSA's retrieval scaffolding moves average human ratings from 0.619 below the expert baseline to statistically indistinguishable from it, with the paired improvement significant at p = 0.019.","PSA roughly doubles the diversity of problem statements: GPT-4o's agent generated 57 unique problems versus 34 for the base model, and Gemini-2.0's agent generated 65 versus 31.","Gemini-2.0-Flash, with or without the PSA pipeline, produced proposals rated above the expert baseline on average, indicating the base model's knowledge sets the quality ceiling.","AI judges cannot replace human blind review in this setting: their scores showed very low variance across all proposals and near-zero correlation with human reviewers, so the paper's main evidence rests on the human panel."],"supporting_citations":[{"why":"Supplies the original four-criteria idea-evaluation rubric that the paper adapts, plus evidence that LLM idea generation lacks diversity and that LLM evaluators disagree with humans.","marker":"Si, Yang, and Hashimoto 2024"},{"why":"Provides the verbalized-confidence technique the PSA uses to score challenge relevance, tractability, and method applicability before sampling.","marker":"Yang, Tsai, and Yamada 2024"},{"why":"Presents the closest prior vision of foundation models for social-impact work, which this paper extends by making scoping fully automated and multi-context.","marker":"Zhao et al. 2024"},{"why":"Represents the iterative, literature-grounded research-agent line that PSA builds on and a reference point for claims about LLM-human agreement in evaluation.","marker":"Lu et al. 2024"},{"why":"Shows that LLM evaluators favor their own generations, which motivates using an independent LLM as the AI judge.","marker":"Panickssery, Bowman, and Feng 2024"},{"why":"Documents inconsistency and bias in LLM evaluators, leading the paper to sample AI evaluations three times and report their low correlation with humans.","marker":"Stureborg, Alikaniotis, and Suhara 2024"},{"why":"Supplies the divergent-convergent creative problem-solving prompting technique that the method-retrieval step adapts against functional fixedness.","marker":"Tian et al. 2024"}],"fun_headline_variants":["LLM scoping agent rivals expert-written AI4SG proposals","Automated scoping matches expert quality for social-good AI","Problem Scoping Agent: LLM proposals score with experts","AI4SG scoping pipeline earns expert-level proposal scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's strongest evidence for expert parity is a blind comparison against expert summaries that were first rewritten by GPT-4o for formatting; the load-bearing premise is that this rewrite preserved the quality and content of the original expert prose.","fun_headline_variants_meta":{"raw":{"variants":["LLM scoping agent rivals expert-written AI4SG proposals","Automated scoping matches expert quality for social-good AI","Problem Scoping Agent: LLM proposals score with experts","AI4SG scoping pipeline earns expert-level proposal scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00019,"raw_usage":{"total_tokens":1315,"prompt_tokens":895,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":364}},"tokens_in":511,"tokens_out":420,"duration_ms":4800,"temperature":1.0,"reasoning_tokens":364,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:36:52.926494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a new panel of blinded reviewers score the original, unrewritten expert summaries alongside PSA proposals with the same four-metric rubric; if the agent no longer matches the baseline, the claimed parity is an artifact of the rewriting step. A smaller test is to score original versus rewritten expert summaries directly; material score differences mean the baseline itself was moved by the rewrite.","supporting_citations":[],"review_version":1}