{"id":"4226d544-7352-43bf-b0ff-9a9b449520ac","arxiv_id":"2607.24352","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.5,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"RAG-augmented local Bielik and PLLuM models produce more terminologically dense Polish legal answers than bare LLMs, yet the study is small, readability-focused, and still reports fundamental legal mistakes.","lead":"Local Polish LLMs plus RAG give denser, more source-tied answers on Polish real-estate law questions than the same models alone, on ordinary hardware. The paper treats that setup as a building block for on-prem regulatory knowledge systems, but the evaluation is a tiny qualitative case study and still surfaces serious legal errors.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline claim (\"significantly improves factual consistency... reducing the risk of unsupported content generation\") is supported only by readability metrics that by design measure no facts — and the paper's own RAG outputs contain unchecked, plausibly fabricated legal citations that would test,","rationale":"The reader already identified the claim–evidence mismatch (four hand-chosen prompts, readability proxies, informal legal review) and the self-reported RAG error, so I agree with their weakest_assumption. My stress test sharpens the same concern into its most falsifiable form: the RAG outputs' specific legal citations are either verifiable or not, and the paper never checks them. This is the load-bearing point because the abstract's claim is precisely about unsupported-content risk — fabricated case numbers in the paper's own Table 2 would not merely weaken the evidence, they would invert the conclusion. The paper deserves credit for transparently reporting one substantive error (§4) and for describing a reproducible on-prem stack (llama.cpp, LM Studio, ChromaDB, Bielik/PLLuM), which supports a narrower contribution. But transparency about one error does not substitute for measuring the quantity the abstract claims to have improved. I keep the verdict at CONDITIONAL rather than REJECT: the architectural/experience-report contribution is real, and the citation-verification test is cheap; if it passes, the paper's claims could be supported with modest additional evaluation. If it fails, the abstract's central sentence must be removed or the paper rejected as stated.","tokens_in":21980,"tokens_out":1739,"duration_ms":61861,"concrete_test":"Verify every specific citation in the Table 2 RAG outputs against primary sources: check case numbers III CZP 83/18 and II CSKP 103/22 in the CBOSA court-rulings database; the interpretation 0114-KDIP1-2.4012.719.2022.2.MC in the KIS tax-interpretation database; and the statutory pinpoint cites (KC Art. 235; UKWH Art. 2 pkt 1; UGN Art. 4 pkt 1; Spatial Planning Act Art. 2 pkt 13; Environmental Protection Act Art. 3 pkt 48) in ISAP. Count how many exist and say what the models claim. If more than ~1 in 5 is fabricated or misattributed, the central \"reduced unsupported content\" claim is contradicted by the paper's own data and must be withdrawn; if all check out, the factual-consistency claim gains real (if still small-n) support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is epistemic: RAG improves \"factual consistency, domain specificity and normative precision\" and reduces hallucination risk (Abstract; §6). But the only quantitative evidence offered (Tables 3–4: Gunning FOG, Jasnopis difficulty score, Logios PLI, parts-of-speech shares) measures text surface properties, and the paper itself concedes these tools \"do not carry out any substantive assessment\" (§2.3). Higher FOG and more \"sophisticated words\" are then reinterpreted in §4 as evidence of normative precision — an inference that conflates lexical density with correctness. The four-prompt, single-topic (real estate) design with no repetition, no inter-rater protocol, and informal author review cannot bear the weight of \"significantly improves.\"\n\nMore damaging, the load-bearing risk is visible in the paper's own data. The RAG outputs are richer precisely in the details most prone to fabrication: PLLuM+RAG (Table 2, RQ4) cites a Supreme Court resolution of 15 Feb 2019 (III CZP 83/18), a judgment of 4 Feb 2022 (II CSKP 103/22), a Ministry of Finance individual interpretation (0114-KDIP1-2.4012.719.2022.2.MC, 15 Jan 2023), and a MRiT statement of 8 Feb 2024; Bielik+RAG (RQ2) attributes the component-parts rule to \"Art. 235 KC\" (the actual provision is Art. 48 KC; Art. 235 concerns a different matter). None of these is verified anywhere in the paper. The author does flag one substantive error (movable vs. real property, §4) and the definitional-conflation flaw, but only via ad-hoc reading — there is no systematic citation-accuracy check. If even some of those pinpoint citations are hallucinated, the paper's own flagship outputs directly contradict the abstract's central claim. This is not an external-consensus objection; it is an internal claim–evidence mismatch with a falsifiable resolution.","agreement_with_reader":"agree"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript proposes local, on-premises use of Polish LLMs—Bielik-1.5B and PLLuM-12B—augmented by retrieval for regulatory knowledge management. It frames the resulting system as a hybrid cognitive-computing architecture in which retrieval supplies controlled knowledge and the LLM performs semantic interpretation. The empirical validation consists of one four-question dialogue about the Polish civil-law definition of real estate, run with and without RAG, followed by readability and syntactic analysis using Jasnopis.pl and Logios.dev. The authors also discuss selected legal errors in the generated answers. They conclude that RAG significantly improves factual consistency, domain specificity, and normative precision, reduces unsupported content, and adds auditability and dynamic knowledge updating.","tokens_in":22319,"tokens_out":7230,"duration_ms":254288,"significance":"Data-sovereign, on-premises retrieval for legal and regulatory work is a timely and practically important problem, and the focus on Polish-language models broadens the literature beyond English-centric systems. Useful strengths include a named consumer-hardware stack, discussion of local legal repositories, full side-by-side generated answers, supplementary Polish originals with English translations, and a candid discussion of one substantive RAG error. If supported by a substantive legal-factuality evaluation and clearer traceability, the work could serve as a valuable proof of concept. At present, however, the evidence supports only an exploratory demonstration, not the broad reliability conclusions.","major_comments":[{"comment":"§2.3 and Tables 3–4 versus the Abstract and §§4, 6: the central claims of improved “factual consistency,” “normative precision,” and reduced unsupported generation are inferred from FOG, PLI, sentence/word length, and parts-of-speech shares. §2.3 correctly states that these tools “do not carry out any substantive assessment.” §4 then treats higher FOG and more sophisticated words as evidence of legal precision, conflating lexical density with correctness. RAG may instead be verbose or copy retrieved material. Four prompts on one concept, one response per condition, no repetitions and no statistical analysis also cannot support “significantly improves.”","section":"§2.3, Tables 3–4, §4, Abstract/§6"},{"comment":"Table 2 contains errors directly bearing on the headline claim. The authors identify PLLuM+RAG’s RQ1 statement that real estate may be movable, but Bielik+RAG’s RQ2 also attributes the component-parts rule to Art. 235 KC, whereas the paper’s own baseline output cites Art. 48; Art. 235 concerns a different subject. PLLuM+RAG’s RQ4 gives detailed Supreme Court, ministerial-interpretation, and MRiT citations without linking any to retrieved documents or checking them. The paper therefore lacks a complete legal-correctness and citation audit while claiming reduced hallucination risk.","section":"Table 2, RQ1–RQ4; §4"},{"comment":"The retrieval architecture is not described consistently. §2.1 names Polish BERT, ChromaDB, LangChain, and vLLM; §2.2 says the LM Studio/Zotero MCP setup “was not a classic vector-based RAG” and used Tavily for web retrieval; §2.3 then describes embedding-based retrieval from ChromaDB. It is unclear which pipeline produced Table 2. Corpus version and scope, chunking, embedding model, top-k, thresholds, reranking, prompts, decoding parameters, seeds, retrieval logs, and code are absent. Consequently the local architecture and its claimed auditability and controlled updating are not reproducibly demonstrated.","section":"§2.1–§2.3"},{"comment":"The cognitive-computing claim is not operationalized or tested. The manuscript never defines measurable criteria by which an LLM becomes a “semantic processing module” within a cognitive-computing infrastructure, nor does it demonstrate integration with control, validation, audit trails, or knowledge-update mechanisms. The experiment compares generated legal prose and readability metrics. Either the CC terminology should be narrowed to a motivation/architectural analogy, or the paper should test concrete properties such as source traceability, controlled repository updates, error containment, and audit logging.","section":"Title, §1, §5, §6"}],"minor_comments":[{"comment":"“Significantly” has a statistical connotation, but no inferential test or repeated sampling is reported. Use “in this exploratory comparison” or provide an appropriate statistical evaluation.","section":"Abstract, §4, §6"},{"comment":"Some prose does not match the table. Bielik+RAG has 34.7% nouns, not “over 40%,” and sophisticated nouns decrease from 10.6% to 6.3% rather than showing the claimed similar increase; PLLuM remains unchanged at 6.0%.","section":"§4, Table 3"},{"comment":"State whether the metrics were computed on all four concatenated answers or separately per RQ, and report per-question values. Explain the Polish adaptation of Gunning FOG and the interpretation/direction of PLI. Table 4 mixes units by reporting “14 lat” and “13 lat.”","section":"Tables 3–4"},{"comment":"Clarify what “better performance” means when Ollama is abandoned, give versions for all components rather than only LM Studio/Zotero, and normalize model identifiers. “MPs” should presumably be “MPS” (Metal Performance Shaders).","section":"§2.1–§2.3"},{"comment":"RQ4 (“Nowhere else?”) is highly context-dependent. Specify whether it was issued in the same conversational session and provide the complete prompt template and session state.","section":"§2.3"},{"comment":"Several references lack complete venue, page, DOI, or publication-status information, including [1], [3], [10], [12], [16], and [18]. Please verify and complete them consistently.","section":"References"},{"comment":"The supplementary tables duplicate the numbers of main Tables 1–2. Label them S1/S2 and state whether the English translations are author-made, machine-assisted, or professionally translated.","section":"Supplementary Materials"},{"comment":"There are recurring grammar and style issues, including “converted to an embeddings,” “there exists laws,” “outputs constructed basing on,” “LLMs is,” and “devoting attention the doctrinal works.”","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"In its present form the manuscript reads more like an exploratory systems demonstration than an empirical validation of the claims in the title and abstract. I recommend major revision rather than rejection because a corrected substantive evaluation could make the setup useful, but publication should not proceed if the authors retain the present headline claims without such an evaluation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a workable experience report on Bielik/PLLuM + local RAG for Polish regulatory Q&A on consumer hardware, not a demonstration that RAG turns LLMs into epistemically reliable cognitive-computing modules. The abstract overclaims what the design can support.\n\nWhat is actually new is narrow and real. You get a concrete on-prem stack (llama.cpp / LM Studio, Chroma, embeddings, MCP/Zotero/Tavily hooks), two Polish models, and a side-by-side on four related real-estate prompts with vs without RAG. The sovereignty motivation is clear, the setup is reproducible in outline, and §4 is unusually honest: the author flags PLLuM+RAG’s movable/real-property confusion and the conflation of civil-law definition with functional uses in special statutes. That self-critique is worth more than most legal-RAG marketing.\n\nWhat it does well stops there. Readability tools (Jasnopis FOG, Logios PLI, POS shares) are used as the only quantitative layer, and the paper itself says they do no substantive assessment. Higher FOG and denser legal diction are then read as “normative precision.” That is a category error. There is no citation-accuracy check, no hallucination metric, no repeated sampling, no second legal rater. Worse, the flagship RAG answers lean hard on pinpoint case numbers, ministry interpretations, and article cites (e.g. Art. 235 KC for component parts; SC and MF references in RQ4) that are never verified in the paper. If those are wrong, the paper’s own outputs undercut the headline claim. The cognitive-computing framing mostly re-labels a standard RAG split (generator vs retrieval memory).\n\nFor whom: practitioners and Polish/local-LLM people who want a worked on-prem recipe and a cautionary failure mode. Not for anyone needing evidence that factual consistency “significantly” improved.\n\nI would send it to peer review as a limited systems/experience paper if the claims are cut down to what four prompts and surface metrics can bear. As written in the abstract, it should not pass without major scoping. Engage if you care about local non-English legal RAG stacks; skip if you need evaluation rigor.","headline":"Useful on-prem Polish RAG case study, but the abstract’s reliability claim outruns four prompts and readability scores—and the paper’s own RAG answers look citation-risky.","tokens_in":20336,"tokens_out":562,"would_cite":false,"duration_ms":19327,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Local LLMs plus RAG become auditable semantic modules for regulatory work, not just free-running generators.","keywords":["Large Language Models (LLMs)","Retrieval-Augmented Generation (RAG)","Cognitive Computing","Regulatory Knowledge Management","Legal Text Analysis","On-Premises AI","Semantic Retrieval","Explainable AI"],"falsifier":"Run the same local Bielik/PLLuM+RAG stack on a larger, blinded set of regulatory questions whose gold answers are fixed statutes and holdings; if source-grounded accuracy and hallucination rate do not improve over the non-RAG baseline under independent legal scoring, the central claim fails.","tokens_in":19944,"feed_emoji":"⚖️","tokens_out":776,"duration_ms":15984,"temperature":0.7,"pith_summary":"This paper asks whether pairing a local large language model with retrieval-augmented generation turns the model from a closed text generator into a reliable piece of cognitive infrastructure for legal and regulatory work. The author runs Polish models (Bielik and PLLuM) on ordinary hardware through Ollama and LM Studio, then compares answers to the same civil-law questions with and without a RAG layer that pulls from curated legal repositories. With RAG, outputs become denser in correct statutory language, cite sources, and can be updated by refreshing the document store rather than retraining the model. The practical stake is on-premises regulatory knowledge management: organizations keep data sovereignty, gain audit trails, and reduce unsupported claims when law changes frequently. The paper therefore argues that such hybrids should be treated as semantic processing modules inside larger cognitive systems, not as standalone chat tools.","feed_headline":"Local LLMs with RAG become auditable legal modules","feed_subtitle":"On-prem Polish models plus retrieval raise normative precision and let rules update without retraining","key_machinery":"Hybrid local RAG architecture: the LLM does semantic interpretation and generation; the retrieval layer (embeddings, vector store or tool-based MCP retrieval over legal corpora) supplies controlled, traceable context so outputs are grounded in current normative sources rather than parametric memory alone.","core_discovery":"Augmenting locally deployed LLMs with RAG significantly improves the factual consistency, domain specificity, and normative precision of generated legal texts while reducing unsupported content, and simultaneously adds auditability and dynamic knowledge update without model retraining; therefore these systems should be regarded as semantic processing modules within cognitive computing infrastructures for regulatory compliance.","pith_inferences":["The same local RAG pattern should transfer to other high-stakes, frequently amended domains (tax, procurement, internal policy) where data residency matters as much as answer quality.","Because the paper itself records a clear category error in one RAG answer, production use still needs an explicit validation or human-in-the-loop gate before any output is treated as normative.","Tool-based retrieval over curated libraries (as with the Zotero/MCP path) may matter as much as pure vector search when legal collections are already expert-organized."],"forward_implications":["Organizations can keep sensitive legal corpora on-premises and still get current, citable answers without shipping data to external model APIs.","Regulatory knowledge bases can be refreshed by re-indexing new acts and judgments; the language model itself need not be retrained.","Generated answers become auditable because retrieved fragments and source metadata travel with the output.","Local consumer-class hardware plus open Polish models is presented as sufficient for practical regulatory assistance pipelines.","LLMs shift role from autonomous knowledge holders to interpretive front-ends inside larger cognitive compliance architectures."],"fun_headline_variants":["Local LLMs with RAG act as auditable regulatory modules","RAG turns on-prem LLMs into semantic compliance processors","Retrieval lifts local LLMs to normative precision without retraining","On-prem Polish LLMs plus RAG add traceable legal knowledge control","Hybrid RAG-LLM architecture yields auditable regulatory modules"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That four hand-chosen prompts about one civil-law definition, judged mainly by readability scores and informal legal review, are enough to show better factual consistency and lower hallucination risk for regulatory work in general.","fun_headline_variants_meta":{"raw":{"variants":["Local LLMs with RAG act as auditable regulatory modules","RAG turns on-prem LLMs into semantic compliance processors","Retrieval lifts local LLMs to normative precision without retraining","On-prem Polish LLMs plus RAG add traceable legal knowledge control","Hybrid RAG-LLM architecture yields auditable regulatory modules"]},"model":"grok-4.5","effort":"low","cost_usd":0.005188,"raw_usage":{"total_tokens":1423,"prompt_tokens":783,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":51884000,"prompt_tokens_details":{"text_tokens":783,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":574,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":783,"tokens_out":66,"duration_ms":9724,"temperature":1.0,"reasoning_tokens":574,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T17:22:14.036011+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same local Bielik/PLLuM+RAG stack on a larger, blinded set of regulatory questions whose gold answers are fixed statutes and holdings; if source-grounded accuracy and hallucination rate do not improve over the non-RAG baseline under independent legal scoring, the central claim fails.","supporting_citations":[],"review_version":1}