{"id":"71c6e429-7551-4c89-ba03-2e9acb58c502","arxiv_id":"2505.10586","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A dynamic RAG pipeline generates situation awareness reports for peacebuilding from GDELT, ACLED, ReliefWeb, and World Bank data, evaluated by NLP metrics, UNDP experts, and LLM judges.","lead":"The paper builds a retrieval-augmented language model system that automatically drafts peacebuilding situation reports from news, conflict event databases, and economic indicators. It introduces a three-level evaluation framework mixing automated NLP metrics, United Nations expert review, and LLM-as-a-judge scoring.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The evaluation lacks a non-RAG baseline and never calibrates automated metrics against expert judgment, so the reported 62–64% scores and the 50% time-saving claim are not attributable to the RAG system.","rationale":"The paper is honest about limitations, and the system design is plausible: dynamic KB construction from four public APIs, top-10 retrieval, and two prompt/LLM variants are concrete and replicable in principle. My concern is not that the system is fraudulent or useless; it is that the evidence as reported does not support the causal claim. The abstract and §5.4 assert that the RAG system reduces analyst burden by ~50% and produces 'coherent, insightful, and actionable reports,' but Section 4 contains no comparison condition that would let a reader attribute those outcomes to the RAG pipeline. A prompt-only LLM might achieve similar or better expert scores on short, structured summaries, especially since the retrieval step is a simple top-10 FAISS search over chunks of news and event data. Furthermore, the three-level evaluation is presented as a validation framework, yet the paper never reports whether Level-1 scores (VERISCORE, SummaC, coherence) agree with Level-2 human judgments; §5.1 even documents large disagreements between LLM judges and humans. Thus the 'framework' provides three parallel opinions, not a chain of validation. A controlled baseline and a calibration analysis are the minimal checks that would settle whether RAG is doing the work claimed. The reader's CONDITIONAL verdict already acknowledges these gaps; my analysis reinforces that condition rather than changing it.","tokens_in":9725,"tokens_out":8405,"duration_ms":86638,"concrete_test":"Run a controlled comparison on the same 15 country/date scenarios: generate 60 additional reports with the identical prompts and LLMs but with retrieval disabled (no evidence injected), producing 120 reports total. Have the two senior UNDP experts score all reports blind on Questions Q1–Q10, and compute (i) a paired significance test (e.g., Wilcoxon or bootstrap) between RAG and no-RAG binary/preference scores, and (ii) the per-report Spearman correlation between Level-1 metrics (VERISCORE, RAG-VERISCORE, SummaC, coherence) and Level-2 human scores. If RAG does not significantly outperform no-RAG, or if Level-1 metrics do not track human judgments, the central quality and time claims lack empirical support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion in §5.4—that the system cuts analyst time by ~50% and produces reports suitable as a first draft—depends on the evaluation actually isolating the RAG system's contribution. It does not. Section 4 enumerates three evaluation levels but no control condition: there is no prompted-only LLM without retrieval, no human-drafted baseline, and no measured drafting/refinement time. The only quantitative human evidence (Table 3) is 62–64% of binary relevance/completeness points, with Cohen's Kappa falling to 0.12–0.17 on preference questions for LLaMA reports. The 'approximately 50%' reduction in §5.4 is an estimate, not a measurement. In addition, the three levels are never calibrated to each other: §5.1 reports GPT-as-a-judge gives self-generated reports 1.0 while human scores are 0.62–0.64, and Claude's preferences diverge from human preferences. Without a baseline and without a demonstrated relationship between Level-1 automated metrics and Level-2 expert judgments, the reported scores do not establish that RAG (rather than the underlying LLM or prompt) produces the quality, nor that the time saving is real.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript describes a dynamic retrieval-augmented generation (RAG) system that automatically drafts situation awareness reports for peacebuilding contexts. It retrieves data from GDELT, ACLED, ReliefWeb, and the World Bank based on a country and date range, embeds the data with MiniLM, retrieves the top-10 evidence items with FAISS, and prompts GPT-4o or LLaMA-3 with either an instruction or a personification prompt to generate a structured report. The authors propose a three-level evaluation: automated NLP metrics (VERISCORE and a modified 'ragve' version, SummaC, politicalBiasBERT, a BERT-based coherence score), expert review by two senior UNDP staff using binary and pairwise questions, and LLM-as-a-judge evaluation. On 15 country/date inputs they generate 60 reports. They report that human experts scored GPT reports 62% and LLaMA reports 64% on binary relevance/completeness, preferred GPT in 76% of pairwise comparisons, and claim the system reduces analyst time by about 50% (from two weeks to one week). They release code and evaluation tools.","tokens_in":9914,"tokens_out":6901,"duration_ms":64958,"significance":"If the reported results were properly supported, the contribution would be useful: a dynamic, publicly sourced RAG pipeline with UNDP expert involvement and open code. The paper's strengths include the use of real expert evaluators from a target organization, the construction of query-specific knowledge bases on demand, and the release of the ragve tool and code. However, the current evaluation does not isolate the RAG contribution, does not establish agreement between automated and human metrics, and does not provide uncertainty quantification. These gaps mean the central empirical claims (report quality and the 50% time saving) are not yet established. The three-level framework is a reasonable idea, but as presented it is not a validated evaluation methodology.","major_comments":[{"comment":"The reported quality scores cannot be attributed to the RAG framework because no baseline is included. There is no condition in which the same prompts are run without retrieval, no human-drafted report baseline, and no ablation varying the number of retrieved evidence items. Without such a control, the 62–64% binary scores and the 'foundation for refinement' claim in Section 5.4 might reflect the underlying LLM and prompt rather than retrieval augmentation.","section":"Section 4, Level 2; Section 5.4"},{"comment":"The claim that the system 'reduces this time by approximately 50%' is an estimate, not a measured outcome. The text states a human analyst currently takes up to two weeks and the system drops this to one week, but no time-motion study, survey, or measurement of drafting/refinement time is reported. The claim should be removed or supported with data, or explicitly labeled as a projected benefit.","section":"Sections 5.2 and 5.4"},{"comment":"The automated Level-1 metrics are never calibrated against the human expert judgments. Table 1 reports average VERISCORE, RAG VERISCORE, SummaC, bias, and coherence scores without variance, confidence intervals, or per-report values, and Section 5.1 shows GPT-as-a-judge assigning its own reports a perfect 1.0 while human scores are around 0.62–0.64. No correlation, regression, or agreement statistic is provided between any Level-1 metric and Level-2 human scores, so the claim that the three-level framework 'enhances reliability' is unsupported.","section":"Table 1 and Section 5.1"},{"comment":"Inter-rater reliability for the preference questions is too low to support the pairwise preference conclusions. For LLaMA-generated reports, Cohen's Kappa for Q8 (completeness) and Q9 (accuracy) is 0.12–0.17, effectively near chance agreement, and the overall preference Cohen's Kappa for LLaMA prompt 2 is 0.31. With only two evaluators and 15 reports per condition, the 76% preference for GPT should be reported with the low agreement explicitly acknowledged, and disagreements should be adjudicated or the preference items discarded.","section":"Table 3, Preference-Based Evaluation"},{"comment":"The system description states that 'Only reports that met acceptable threshold values in Level 1 evaluation were forwarded to the next stage,' but no threshold values are defined and Table 3 reports evaluation results for all four model-prompt conditions. It is therefore unclear whether any reports were filtered and what the thresholds were. The thresholds and the number of reports passing or failing at each threshold should be stated, or the sentence should be corrected to say that no filtering was applied.","section":"Section 4, Level 1"},{"comment":"The LLM-as-a-judge component uses the same model families that generated the reports (GPT-4o and LLaMA 3), and the paper reports that GPT assigns its own reports a perfect 1.0. The authors acknowledge 'possible overconfidence or evaluation bias,' but the framework still includes these scores as evidence of robustness. The self-evaluation scores should be either removed from the headline results, corrected with a debiased judge, or replaced with a model that was not used for generation and whose scores are calibrated to human judgments.","section":"Section 4, Level 3 and Section 5.1"}],"minor_comments":[{"comment":"The GitHub URL for ragve is given as 'withheld-for-anonymity'; this placeholder should be replaced with the actual repository URL or removed if the paper is under anonymized review.","section":"Section 4, Level 1, footnote 11"},{"comment":"The literature review contains an unresolved citation placeholder '[?]' immediately after [Joshi et al., 2024].","section":"Section 2"},{"comment":"The text says 'Horns of Africa'; the correct region name is 'Horn of Africa'.","section":"Section 4, Test Sample Selection"},{"comment":"The row 'Avg. Max Score (Binary Questions)' reports 0.62, 0.64, 0.60, 0.63 for the four conditions, but Section 5 states GPT reports achieved 62% and LLaMA reports scored 64%; the relationship between the row and the quoted percentages is unclear, and averaging the LLaMA columns gives 0.615, not 0.64.","section":"Table 3 and Section 5"},{"comment":"The 'Politically Center Confidence' is 0.99 for all four conditions, a ceiling effect that is never discussed; the bias metric therefore appears non-diagnostic and should either be analyzed per report or removed from the evaluation.","section":"Table 1, Objectivity/Bias"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a lightly revised preprint: the GitHub URL is anonymized and there is an unresolved citation placeholder, which suggest it may have been submitted before careful proofreading. The claims of being 'first' in peacebuilding RAG are strong and could benefit from a more extensive literature check, but that is not the main obstacle. The core problem is that the evaluation design lacks a non-RAG baseline, does not calibrate automated metrics to human judgments, and includes an LLM judge that self-scores perfectly. These are fixable in revision: adding a no-retrieval condition, reporting variance and calibration, and removing or debiasing the self-judge would make the central claims testable. The involvement of two UNDP experts is a genuine asset and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is the first documented RAG application to peacebuilding report generation, and it gets real UNDP experts into the loop; but the evaluation design doesn't isolate the RAG system's contribution, and the headline time-saving claim is an estimate rather than a measurement.\n\nWhat's new: the domain application, the dynamic knowledge base built from four public APIs (GDELT, ACLED, ReliefWeb, World Bank), and ragve, a modified VERISCORE that checks grounding against the retrieved knowledge base instead of a web search. That tool is a genuine small contribution, worth citing if you work on RAG verification. The three-level evaluation design (automated metrics, human experts, LLM judges) is sensible on paper, and the authors are honest about the human-in-the-loop requirement and regional coverage gaps.\n\nWhere it goes soft: there is no baseline. No prompted-only LLM without retrieval, no human-drafted report, no measured drafting or refinement time. So the 62-64% human binary scores and the \"~50% time reduction\" in §5.4 cannot be attributed to the RAG mechanism rather than the underlying LLM or prompt. There are no error bars or statistical tests anywhere; the sample is 15 inputs and 2 experts; the human Cohen's Kappa on preference questions for LLaMA reports drops to 0.12-0.17; Level 1 acceptance thresholds are never stated; and the authors acknowledge GPT-as-a-judge giving its own reports a perfect 1.0 but never quantify the self-preference confound. The anonymized GitHub URL also blocks reproducibility claims.\n\nThe stress-test note is fair: the three evaluation levels are never calibrated to each other, so the automated metrics don't back the human scores, and the 50% figure is an estimate. That said, this doesn't sink the paper. The authors explicitly frame the system as a first-draft generator with mandatory human review, and the limitations section is candid. The core feasibility claim—that a RAG pipeline can produce structured, partially credible situation reports from public data—is plausibly demonstrated, though not rigorously proven.\n\nWho this is for: applied NLP researchers, peacebuilding and humanitarian tech folks, and RAG practitioners wanting a real-world case study with expert feedback. It deserves a serious referee round, but the revision bar should be high: add a non-RAG baseline, report variance, state thresholds, release the code, and either measure time savings or drop the claim.\n\nRecommendation: accept for peer review, not desk reject, with clear guidance for a major revision.","headline":"First RAG-for-peacebuilding paper with real UNDP experts, but the evaluation never isolates the RAG contribution and the 50% time-saving is asserted, not measured.","tokens_in":10522,"tokens_out":2489,"would_cite":false,"duration_ms":24252,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a retrieval-augmented generation system can turn public conflict, news, and economic data into first-draft peacebuilding reports, cutting analyst time by roughly half.","keywords":["retrieval-augmented generation","situation awareness","peacebuilding reports","conflict monitoring","expert evaluation","LLM-as-a-judge","factual consistency","open data"],"falsifier":"Take twelve country-periods balanced between high- and low-coverage regions, have analysts list every fact a complete situation report should contain, and count how many appear in the RAG report; if low-coverage regions systematically miss a large share of reference facts, the knowledge-base claim is shown to be coverage-limited. A second check: compute the correlation between the Level-1 automated metrics and the Level-2 expert binary scores across all 60 reports; near-zero correlation would falsify the claim that the three-level framework is a coherent measure of report quality.","tokens_in":9489,"feed_emoji":"🕊️","tokens_out":8524,"duration_ms":81529,"temperature":0.7,"pith_summary":"The paper tries to establish that automated situation-awareness reports for peacebuilding are achievable with a retrieval-augmented generation (RAG) system that assembles, on demand, a knowledge base from public news, conflict-event, humanitarian, and economic data and then produces a structured, source-cited first draft. A sympathetic reader would take the central message to be that this first draft is good enough to cut a human analyst's work from about two weeks to roughly one week, while leaving final judgment to experts. The authors support this with a three-level evaluation: automated NLP metrics for factuality, consistency, bias, and coherence; binary relevance and completeness questions plus pairwise preferences answered by two senior experts from the partner humanitarian agency; and an LLM-as-a-judge pass using the same questionnaire. Their reported results show moderate human inter-annotator agreement, a 76% human preference for GPT-generated reports over LLaMA-generated ones in pairwise choices, and LLM judges scoring more generously than humans, which the paper takes as evidence that expert review remains mandatory.","feed_headline":"RAG pipeline drafts peacebuilding reports in half the time","feed_subtitle":"The system pulls public news, conflict events, and economic data into evidence-backed first drafts for human analysts.","key_machinery":"The load-bearing mechanism is the dynamic retrieval pipeline: a user-specified country and reporting window trigger queries to four public data APIs, the fetched documents are cleaned, numeric values are wrapped in textual templates, and a lightweight sentence encoder embeds everything into a query-specific vector store. A similarity search over that store selects the top ten evidence passages, which are pasted into a prompt that instructs the LLM to produce a structured report with source citations. The second half of the machinery is the three-level evaluation stack: automated metrics (VERISCORE for factuality, a modified ragve variant that verifies grounding in the retrieved knowledge base, SummaC for consistency, a political-bias classifier, and a BERT-based coherence score), expert evaluation through binary relevance and completeness questions plus pairwise preferences, and LLM-as-a-judge scoring with the same questionnaire. The ragve modification is what ties the evaluation to the RAG claim, because it checks that claims are grounded in the retrieved evidence rather than in model memory.","core_discovery":"On the paper's own terms, the discovery is that a dynamic RAG framework, the first applied to peacebuilding, can generate reports that are coherent, insightful, and actionable enough to serve as a foundation for expert refinement. The system takes a country and a date range, queries four public data sources, converts numeric indicators into text, embeds the collected documents, retrieves the top ten most similar passages for the query 'Conflict and social unrest issues in {country}', and prompts an LLM to write a report with fixed sections and explicit source citations. The paper reports factuality scores averaging 0.73-0.91 on VERISCORE and 0.60-0.69 on a knowledge-base-grounded variant, expert binary scores around 62-64% of the maximum, human preference for GPT-generated reports in 76% of pairwise comparisons, and a claimed time saving of approximately 50% because the generated report becomes the base for review and refinement.","pith_inferences":["A natural experiment the paper does not run is varying the number of retrieved passages (for example, five versus twenty) and measuring expert preference, which would test whether the top-ten budget is optimal or whether more evidence improves completeness at the cost of noise.","The paper never correlates its Level-1 automated scores with the Level-2 human binary judgments; computing that correlation across the 60 reports would show whether the three-level framework measures one coherent notion of quality or three different things.","The regional variation in performance implies a testable prediction: for a fixed country, RAG report quality should track the volume and recency of international media coverage, which could be quantified from the news-event API and regressed against expert scores.","The divergence among LLM judges suggests that which model wrote the report and which model judges it jointly determine evaluation outcomes, so any future deployment should report judge identity alongside quality scores."],"forward_implications":["First-draft situation reports can be generated in near real time from open data, cutting the manual analysis cycle from roughly two weeks to one and letting the same analyst team cover more countries or more frequent updates.","Report quality is bounded by the retrieved evidence: in low-coverage regions such as the Horn of Africa, performance degrades, so operational deployments must supplement the four sources or state coverage limits explicitly.","LLM-as-a-judge cannot replace human evaluation; the paper's data show self-preference (GPT scoring its own reports at 1.0) and cross-model disagreement, so automated judges need calibration against human anchors before use.","The three-level framework gives other teams a reusable template for validating RAG output in high-stakes settings, with the modified factuality checker released as an open tool.","The system is positioned as an augmentation rather than a replacement, since the authors state that human review remains mandatory before reports reach stakeholders."],"supporting_citations":[{"why":"supplies the VERISCORE factuality metric that the paper adapts into the knowledge-base-grounded ragve check.","marker":"[Song et al., 2024]"},{"why":"supplies the SummaC consistency measure used to check that reports preserve key points from the retrieved evidence.","marker":"[Laban et al., 2022]"},{"why":"provides the vector similarity search library used to select the top ten evidence passages for each query.","marker":"[Douze et al., 2024]"},{"why":"defines the conflict-event dataset that is one of the four public data inputs to the RAG pipeline.","marker":"[Raleigh et al., 2010]"},{"why":"provides the political-bias classifier used in Level 1 to measure objectivity of generated reports.","marker":"[Baly et al., 2020]"},{"why":"defines the inter-annotator agreement statistic used to quantify consistency between the two human expert evaluators.","marker":"[Cohen, 1960]"}],"fun_headline_variants":["RAG brings automated situation awareness to peacebuilding","RAG pipeline drafts peacebuilding reports from live data","Automated RAG framework for faster peacebuilding analysis","RAG generates evidence-backed peacebuilding reports automatically"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The system's usefulness rests on the assumption that the top ten passages retrieved from the four public data sources contain the balanced facts a good peacebuilding report needs; where news coverage is thin, the report inherits those gaps and the claimed time saving buys a weaker draft.","fun_headline_variants_meta":{"raw":{"variants":["RAG brings automated situation awareness to peacebuilding","RAG pipeline drafts peacebuilding reports from live data","Automated RAG framework for faster peacebuilding analysis","RAG generates evidence-backed peacebuilding reports automatically"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1606,"prompt_tokens":964,"completion_tokens":642,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":580,"completion_tokens_details":{"reasoning_tokens":581}},"tokens_in":580,"tokens_out":642,"duration_ms":6002,"temperature":1.0,"reasoning_tokens":581,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:27:44.409780+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take twelve country-periods balanced between high- and low-coverage regions, have analysts list every fact a complete situation report should contain, and count how many appear in the RAG report; if low-coverage regions systematically miss a large share of reference facts, the knowledge-base claim is shown to be coverage-limited. A second check: compute the correlation between the Level-1 automated metrics and the Level-2 expert binary scores across all 60 reports; near-zero correlation would falsify the claim that the three-level framework is a coherent measure of report quality.","supporting_citations":[{"cited_title":"Summac: Re- visiting nli-based models for inconsistency detection in summarization","cited_arxiv_id":null,"evidence_quote":"supplies the SummaC consistency measure used to check that reports preserve key points from the retrieved evidence."},{"cited_title":"The faiss library","cited_arxiv_id":null,"evidence_quote":"provides the vector similarity search library used to select the top ten evidence passages for each query."},{"cited_title":"Linke, H ˚avard Hegre, and Joakim Karlsen","cited_arxiv_id":null,"evidence_quote":"defines the conflict-event dataset that is one of the four public data inputs to the RAG pipeline."},{"cited_title":"We can detect your bias: Predicting the political ideology of news articles","cited_arxiv_id":null,"evidence_quote":"provides the political-bias classifier used in Level 1 to measure objectivity of generated reports."}],"review_version":1}