{"id":"c4360584-33c4-4264-aff8-56b119420d6a","arxiv_id":"2511.10354","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ATR4CH is a replicable five-step methodology for LLM-based knowledge extraction from cultural heritage documents that combines annotation models and ontological frameworks, achieving F1 scores of 0.96-0.99 for metadata, 0.7-0.8 for entities, and 0.62 G-EVAL for discourse on Wikipedia articles about ","lead":"The paper introduces ATR4CH, a five-step methodology that uses large language models guided by ontologies to convert cultural heritage texts into structured knowledge graphs. A smart generalist might read it to understand how AI can help museums and researchers turn messy debates about artifacts and authenticity into searchable, queryable data.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Wikipedia-only evaluation leaves generalizability to primary CH sources untested","rationale":"The reader's weakest assumption correctly isolates the evaluation scope and human-oversight requirement as the point where the replicability claim is least secure. The concrete test directly measures whether the reported performance holds on primary texts, which is the minimal empirical check needed to support or weaken the strongest claim of a domain-adaptable framework. No deeper internal inconsistency in the five-step description or metric definitions was evident from the reported results.","tokens_in":1806,"tokens_out":450,"duration_ms":44442,"concrete_test":"Apply the full ATR4CH pipeline (annotation schema, sequential Claude Sonnet 3.7 / Llama 3.3 70B / GPT-4o-mini extraction, and ontology integration) to a new corpus of 20 primary-source documents on authenticity debates (e.g., open digitized excavation reports or letters from Europeana or a museum archive). Compute the identical F1 scores for hypothesis extraction and G-EVAL for discourse representation; a drop of more than 15 points relative to the Wikipedia baseline would show the framework does not yet generalize beyond encyclopedic text.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ATR4CH supplies the first replicable five-step methodology (foundational analysis through comprehensive evaluation) for ontology-guided LLM extraction of nuanced scholarly debates into KGs. For this to hold, the pipeline must produce reliable hypothesis and discourse structures from representative Cultural Heritage texts. All reported metrics (0.7-0.8 entity F1, 0.65-0.75 hypothesis F1, 0.95-0.97 evidence F1, 0.62 G-EVAL discourse) derive exclusively from Wikipedia articles on disputed items. These are secondary, consensus-oriented summaries with explicit structure and lower ambiguity than primary sources such as excavation reports, correspondence, or scholarly monographs, which contain denser domain terminology, implicit reasoning, and contested interpretations. The paper itself notes that human post-processing oversight remains necessary; this combination makes the assumption that ontology prompting plus LLM extraction will transfer to authentic CH discourse the least secure element of the replicability claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces ATR4CH, a five-step methodology (foundational analysis, annotation schema development, pipeline architecture, integration refinement, and comprehensive evaluation) that combines LLMs with Cultural Heritage ontologies to extract structured knowledge graphs from texts, with a focus on scholarly debates such as authenticity assessments. It validates the approach in a case study on Wikipedia articles about disputed items using Claude Sonnet 3.7, Llama 3.3 70B, and GPT-4o-mini, reporting F1 scores of 0.96-0.99 (metadata), 0.7-0.8 (entity recognition), 0.65-0.75 (hypothesis), 0.95-0.97 (evidence), and 0.62 G-EVAL (discourse).","tokens_in":2046,"tokens_out":563,"duration_ms":30297,"significance":"If the central claims hold, ATR4CH would supply the first replicable framework for CH institutions to convert unstructured discourse into queryable KGs, enabling metadata enrichment and knowledge discovery. The manuscript earns credit for its concrete empirical metrics across three LLMs (including competitive results from smaller models), explicit acknowledgment of the need for human post-processing oversight, and presentation of a sequential pipeline that integrates annotation models with ontological frameworks.","major_comments":[{"comment":"Case Study / Evaluation section: All reported metrics (0.65-0.75 F1 for hypothesis extraction, 0.62 G-EVAL for discourse representation) derive exclusively from Wikipedia articles on disputed items. These are secondary, consensus-oriented summaries with explicit structure and lower ambiguity than primary CH sources such as excavation reports or scholarly monographs; this scope limitation directly weakens the replicability claim that ATR4CH supplies an adaptable framework across CH domains.","section":"Case Study / Evaluation"}],"minor_comments":[{"comment":"Abstract: The findings paragraph reports '0.95-0.97 for evidence extraction' without specifying the metric; this should be clarified as F1 to maintain consistency with the other reported scores.","section":"Abstract"},{"comment":"Findings: A summary table comparing F1 and G-EVAL scores across the three LLMs for each extraction task (metadata, entities, hypotheses, evidence, discourse) would improve readability and allow direct comparison of model performance.","section":"Findings"},{"comment":"Research Limitations: The statement that 'human oversight is necessary during post-processing' could be expanded with concrete examples of the types of errors or nuances that require intervention.","section":"Research Limitations"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive and detailed review, as well as for recognizing the potential significance of ATR4CH and the value of our empirical results across multiple LLMs. We address the major comment below.","responses":[{"response":"We agree that the evaluation is confined to Wikipedia articles on disputed items, which are secondary sources with relatively explicit structure. The manuscript already states this limitation explicitly in the Research Limitations section: 'The produced KG is limited to Wikipedia articles.' Wikipedia articles were chosen for the case study because they contain accessible, well-documented examples of scholarly debates on authenticity assessments, enabling direct testing of the pipeline's ability to extract hypotheses, evidence, and discourse relations. The ATR4CH methodology itself consists of five general steps (foundational analysis, annotation schema development, pipeline architecture, integration refinement, and comprehensive evaluation) that are intended to be repeatable and adaptable to other CH texts and ontologies. The replicability claim therefore refers primarily to the systematic process rather than to the specific numerical results generalizing unchanged to primary sources. Nevertheless, the referee's point is well taken: stronger evidence of adaptability would require evaluation on primary documents such as excavation reports. We will revise the manuscript to (a) more explicitly frame the current case study as a proof-of-concept demonstration, (b) add a dedicated subsection discussing concrete adaptations needed for less-structured primary sources, and (c) moderate the language around cross-domain adaptability to better reflect the present scope.","revision_made":"partial","referee_comment":"[Case Study / Evaluation] Case Study / Evaluation section: All reported metrics (0.65-0.75 F1 for hypothesis extraction, 0.62 G-EVAL for discourse representation) derive exclusively from Wikipedia articles on disputed items. These are secondary, consensus-oriented summaries with explicit structure and lower ambiguity than primary CH sources such as excavation reports or scholarly monographs; this scope limitation directly weakens the replicability claim that ATR4CH supplies an adaptable framework across CH domains."}],"tokens_in":1533,"tokens_out":423,"duration_ms":36851,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper gives a concrete five-step method called ATR4CH for pulling structured knowledge out of cultural heritage texts with LLMs, focused on scholarly debates like authenticity questions. It combines annotation, ontologies, and model prompting in a replicable sequence and reports solid numbers on a case study.","headline":"ATR4CH lays out a usable five-step LLM pipeline for turning cultural heritage texts into KGs, but the Wikipedia-only tests leave the harder cases unproven.","tokens_in":2534,"tokens_out":141,"would_cite":false,"duration_ms":21139,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"ATR4CH LLM-ontology KG pipeline for CH discourse extraction lies outside RS forcing chain","alignment":"orthogonal","rationale":"The paper's central machinery is a five-step iterative methodology (foundational analysis, MWA annotation, pipeline architecture, integration, evaluation) that coordinates LLMs (GliNER + Claude/Llama/GPT) with SEBI ontology and Wikidata for extracting metadata, Cognizers, evidence, hypotheses and RDF-star claims from Wikipedia articles on authenticity debates. No component invokes recognition cost J(x), golden-ratio identities, φ-ladder spacings, 8-tick periodicity, ratio-symmetric forcing, or parameter-free derivation from a single distinction. The domain (cs.CL / Digital Humanities) is one on which the RS framework expresses no structural opinion.","tokens_in":57484,"confidence":"high","tokens_out":177,"duration_ms":9504,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ATR4CH is a five-step methodology that guides large language models with cultural heritage ontologies to turn texts on scholarly debates into structured knowledge graphs.","keywords":["knowledge graphs","large language models","cultural heritage","ontologies","text extraction","scholarly debates","ATR4CH","information extraction"],"falsifier":"Applying the full ATR4CH pipeline to a collection of primary museum or archive documents on disputed artifacts and finding that expert review identifies more than 30 percent mismatch in extracted hypotheses or discourse structures compared with the Wikipedia results.","tokens_in":2714,"feed_emoji":"🏛️","tokens_out":786,"duration_ms":44218,"temperature":0.7,"pith_summary":"The paper presents ATR4CH as a systematic process for extracting knowledge from cultural heritage documents and converting it into queryable knowledge graphs. It outlines five iterative steps that develop annotation schemas and integrate them with ontological frameworks to direct LLM outputs toward metadata, entities, hypotheses, evidence, and discourse relations. The approach is tested on Wikipedia articles about authenticity debates for artifacts and documents, producing high extraction accuracy even with smaller models. A reader would care because cultural heritage collections contain extensive textual material that stays difficult to search or analyze until it exists in structured form. The work claims to supply the first coordinated framework linking LLMs to established cultural heritage ontologies for this conversion.","feed_headline":"Method turns cultural heritage debates into knowledge graphs","feed_subtitle":"ATR4CH uses five iterative steps and ontology guidance to extract metadata, entities, hypotheses and evidence from Wikipedia texts with F1s ","key_machinery":"ATR4CH, the five-step adaptive text-to-RDF methodology that combines annotation models, ontological frameworks, and LLM-based extraction to convert unstructured cultural heritage texts into RDF knowledge graphs.","core_discovery":"ATR4CH supplies the first systematic methodology for coordinating LLM-based extraction with Cultural Heritage ontologies by progressing through foundational analysis, annotation schema development, pipeline architecture, integration refinement, and comprehensive evaluation. In the authenticity assessment case study on Wikipedia articles, the sequential pipeline with Claude Sonnet 3.7, Llama 3.3 70B, and GPT-4o-mini reached F1 scores of 0.96-0.99 for metadata extraction, 0.7-0.8 for entity recognition, 0.65-0.75 for hypothesis extraction, 0.95-0.97 for evidence extraction, and 0.62 G-EVAL for discourse representation, with smaller models performing competitively.","pith_inferences":["Testing ATR4CH on primary sources such as museum catalogs or historical records would show whether ontology guidance holds for text styles outside Wikipedia.","Linking the resulting graphs to existing digital heritage platforms could enable ongoing updates as new debates appear in the literature.","The same coordinated LLM-ontology pattern might transfer to other areas of contested knowledge such as legal opinions or scientific controversies.","Wider adoption could lower the entry cost for smaller institutions to contribute to linked open data efforts in the cultural heritage sector."],"forward_implications":["Cultural heritage institutions can convert textual knowledge into queryable knowledge graphs without prohibitive manual effort.","Automated metadata enrichment and knowledge discovery become practical for large document collections.","Smaller language models support cost-effective deployment while maintaining competitive extraction performance.","The framework adapts across different cultural heritage domains and varying levels of institutional resources.","Post-processing human oversight remains part of the workflow to finalize the knowledge graphs."],"fun_headline_variants":["ATR4CH builds knowledge graphs from cultural heritage texts","LLMs turn cultural heritage texts into queryable knowledge graphs","Five-step approach creates knowledge graphs from heritage documents","Extraction pipeline applies ontologies to cultural heritage debates"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That large language model outputs guided by ontologies can reliably capture nuanced scholarly debates and discourse structures when the tests use only Wikipedia articles and rely on post-processing human oversight.","fun_headline_variants_meta":{"raw":{"variants":["ATR4CH builds knowledge graphs from cultural heritage texts","LLMs turn cultural heritage texts into queryable knowledge graphs","Five-step approach creates knowledge graphs from heritage documents","Extraction pipeline applies ontologies to cultural heritage debates"]},"model":"grok-4.3","cost_usd":0.008631,"raw_usage":{"total_tokens":3980,"prompt_tokens":840,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":86312000,"prompt_tokens_details":{"text_tokens":840,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3081,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":840,"tokens_out":59,"duration_ms":23586,"temperature":1.0,"reasoning_tokens":3081,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-17T22:21:43.821382+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Applying the full ATR4CH pipeline to a collection of primary museum or archive documents on disputed artifacts and finding that expert review identifies more than 30 percent mismatch in extracted hypotheses or discourse structures compared with the Wikipedia results.","supporting_citations":[],"review_version":1}