{"id":"b2fe436b-fe2b-4bab-9337-52a5aeb84988","arxiv_id":"2604.11724","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"OCR for basic letters in Old Church Slavonic can reach 2-3% error rate with modern systems, and a new image-only pipeline of glyph extraction, clustering, and pairwise comparison can produce stemmata for historical manuscripts.","lead":"This paper compares OCR systems on 18th-century handwritten Old Church Slavonic and introduces a purely visual pipeline for stemma reconstruction using glyph extraction, clustering, and statistical comparison on manuscript images. A smart generalist might read it to see how AI can help digitize and map relationships in historical texts without full manual transcription.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Visual glyph clustering and pairwise image statistics may reconstruct scribal style or script family rather than textual copying relationships.","rationale":"The reader's weakest assumption matches the load-bearing gap exactly. The OCR experiments are more self-contained (CER numbers on a fixed character set), but the stemma claim is the paper's forward-looking contribution and remains unanchored by any external validation. No change to the UNVERDICTED/LOW verdict is warranted.","tokens_in":1826,"tokens_out":338,"duration_ms":41186,"concrete_test":"For the 14th–16th c. Church Slavonic Gospel of Mark corpus, obtain a philologist-provided reference stemma and compute Robinson-Foulds distance (or normalized quartet distance) between the visual method's tree and the reference; if the distance exceeds that of a random tree or a simple script-style clustering baseline, the visual pipeline does not recover copying relationships.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's novel contribution is a purely visual stemma pipeline (glyph extraction, clustering, statistical pairwise distances) applied to two small corpora. Traditional stemmatology reconstructs descent from shared textual variants (errors, omissions, substitutions). Visual shape similarity can arise from shared scribal training, regional hand, or exemplar layout without implying direct copying. The abstract states only that 'basic functioning of the method can be demonstrated'; no quantitative comparison to expert stemmata, no textual variant baseline, and no ablation isolating visual vs. textual signal is described. If the distance matrix primarily encodes non-genetic visual features, the downstream stemma does not support the claimed historical reconstruction.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript compares a range of OCR systems (classical, ML-based, and LLMs such as GPT-5 and Gemini-3-flash) on approximately 6,000 characters from 18th-century handwritten Old Church Slavonic manuscripts, evaluates post-processing and agentic architectures, and reports that CER for basic letters can reach 2-3% while noting persistent issues with diacritics. It then introduces a purely visual stemma reconstruction pipeline consisting of automated glyph extraction, clustering, pairwise statistical comparisons to produce a distance matrix, and applies the method to two small corpora (14th-16th century Church Slavonic Gospel of Mark and 14th-15th century French Roman de la Rose), claiming that basic functioning of the method can be demonstrated.","tokens_in":1956,"tokens_out":602,"duration_ms":35538,"significance":"If the visual stemma pipeline can be shown to recover genealogical relationships rather than merely scribal style or script family, it would constitute a novel contribution to computational stemmatology by operating independently of textual variants. The OCR experiments provide practical insights into current LLM capabilities for historical scripts. However, the current lack of validation metrics, ground-truth comparisons, or controls for confounding visual factors substantially limits the demonstrated impact.","major_comments":[{"comment":"In the stemma reconstruction section, the pipeline is applied to two small corpora and 'basic functioning' is claimed, yet no quantitative validation metrics (e.g., agreement with expert stemmata), error analysis, or controls for non-genetic visual similarity (shared scribal training, regional hand, or layout) are reported. This is load-bearing for the central claim that the distance matrix reconstructs historical copying relationships.","section":"stemma reconstruction pipeline"},{"comment":"No ablation isolating visual glyph statistics from textual content, and no baseline comparison against traditional variant-based stemmatology, is presented. Without such tests it remains unclear whether the pairwise distances encode descent or merely visual similarity, directly affecting the validity of the downstream stemma.","section":"stemma reconstruction pipeline"}],"minor_comments":[{"comment":"The abstract states that 'more than 10 CS OCR-systems among which 2 LLMs (GPT5 and Gemini3-flash) are being compared' but does not list the full set of systems or the precise conditions under which the 2-3% CER is achieved; this should be clarified with a table of results.","section":"abstract"},{"comment":"Several sentences contain awkward or passive phrasing (e.g., 'Focussing on basic letter correctness, more than 10 CS OCR-systems ... are being compared' and 'With new technology elaborated, experiments suggest...'). Consider revising for directness and readability.","section":"abstract and introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed feedback on our manuscript. The primary concerns relate to the validation and interpretation of the visual stemma reconstruction pipeline. We address these points directly below while maintaining that the work offers a proof-of-concept demonstration of a text-independent approach, consistent with the limited scope and small corpora described.","responses":[{"response":"We agree that quantitative validation metrics, error analysis, and explicit controls for non-genetic factors would strengthen the presentation. For the small historical corpora examined, however, no consensus expert stemmata exist for direct quantitative comparison, and constructing reliable ground truth would require extensive philological expertise outside the computational scope of this study. The claim of 'basic functioning' is supported by the pipeline producing distance matrices and clusterings that align qualitatively with known manuscript dates and provenances. We will revise the manuscript to include a more detailed discussion of these limitations, potential confounders such as shared scribal training or layout, and directions for future controlled validation.","revision_made":"partial","referee_comment":"In the stemma reconstruction section, the pipeline is applied to two small corpora and 'basic functioning' is claimed, yet no quantitative validation metrics (e.g., agreement with expert stemmata), error analysis, or controls for non-genetic visual similarity (shared scribal training, regional hand, or layout) are reported. This is load-bearing for the central claim that the distance matrix reconstructs historical copying relationships."},{"response":"The pipeline is constructed to use only visual glyph features extracted from images, with no textual content or OCR output entering the clustering or distance matrix computation; an ablation separating visual statistics from text is therefore satisfied by design. A direct baseline against traditional variant-based methods is not provided because those approaches require accurate transcriptions, whose challenges for these scripts are precisely the subject of the OCR experiments in the first half of the paper. The visual method is offered as an alternative precisely when transcription is unreliable. We will revise the text to clarify this motivation and the interpretive scope of the resulting distance matrix.","revision_made":"partial","referee_comment":"No ablation isolating visual glyph statistics from textual content, and no baseline comparison against traditional variant-based stemmatology, is presented. Without such tests it remains unclear whether the pairwise distances encode descent or merely visual similarity, directly affecting the validity of the downstream stemma."}],"tokens_in":1533,"tokens_out":508,"duration_ms":50244,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper splits into an OCR survey on 18th-century handwritten Church Slavonic and a new image-only stemma method. On the OCR side they test classical, ML, and LLM systems across roughly 6000 characters, then add LLM post-processing and agentic setups. They report basic-letter CER down to 2-3 percent, with diacritics still problematic. That section is straightforward and gives anyone working with this script a current baseline of what is feasible now.","headline":"The OCR comparisons give a useful snapshot of current tools for Church Slavonic, but the visual stemma pipeline only shows it runs on small sets without evidence it recovers actual copying history.","tokens_in":2403,"tokens_out":178,"would_cite":false,"duration_ms":27753,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A pipeline that extracts glyphs from manuscript images, clusters them, and builds distance-based stemmas can reconstruct copying relationships without any text transcription.","keywords":["OCR","stemmatics","Old Church Slavonic","visual analysis","manuscripts","glyph clustering","distance matrix","image processing"],"falsifier":"Running the visual pipeline on a corpus whose copying relationships are already known from textual collation and finding that the resulting stemma differs substantially from the established tree.","tokens_in":2723,"feed_emoji":"📜","tokens_out":728,"duration_ms":53125,"temperature":0.7,"pith_summary":"The paper first benchmarks multiple OCR approaches on late handwritten Old Church Slavonic manuscripts and reports that basic-letter character error rates can reach 2-3 percent with current systems including large language models, while diacritics remain difficult. It then introduces and tests a purely visual stemmatology method that automatically extracts letter shapes from page images, groups similar glyphs, compares pairs statistically, and converts the resulting distances into a family tree. If the method works, researchers could trace manuscript lineages directly from scans of texts whose language or script makes full transcription expensive or error-prone. The demonstration on two small corpora, one Church Slavonic and one French, shows that the pipeline produces plausible stemmas from image data alone.","feed_headline":"Visual glyph shapes alone can build manuscript stemmas","feed_subtitle":"An image-processing pipeline extracts letters, compares them statistically, and produces family trees for Church Slavonic and French texts.","key_machinery":"The visual glyph extraction, clustering, and pairwise statistical comparison pipeline that produces a distance matrix and stemma directly from page images.","core_discovery":"The author presents a complete image-only pipeline for stemma construction: visual glyph extraction from manuscript pages, unsupervised clustering of letter forms, pairwise statistical comparison to form a distance matrix, and conversion of that matrix into a stemma diagram. When applied to a small set of 14th-16th century Church Slavonic Gospel of Mark copies and a set of 14th-15th century Roman de la Rose manuscripts, the pipeline produces tree structures that the author treats as evidence of basic functionality. The claim is that visual shape information alone can serve as a workable proxy for historical copying relationships.","pith_inferences":["If the method scales beyond the two small test sets, it could reduce the bottleneck of manual transcription in digital humanities projects that aim to map manuscript traditions.","The approach might be combined with existing textual stemmatology tools so that visual distances serve as one data layer among others rather than the sole input.","Limitations observed with diacritics in the OCR section suggest that glyph features involving fine marks may need special handling in any larger deployment of the visual pipeline."],"forward_implications":["OCR output can feed into stemmatology even when full transcription is imperfect.","Stemma reconstruction becomes feasible for manuscripts whose scripts are hard to transcribe reliably.","The distance matrix from glyph comparisons offers a quantitative starting point that can later be refined with textual or codicological data.","The same visual approach can be tested on additional small corpora to check consistency across languages and periods."],"fun_headline_variants":["Visual glyphs build manuscript stemmas","Glyph shapes alone build manuscript stemmas","Visual shapes reconstruct manuscript stemmas","Image glyphs form manuscript stemmas"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That letter shapes captured in images, without any textual content or expert judgment, contain enough information to recover accurate historical copying relationships on the tested corpora.","fun_headline_variants_meta":{"raw":{"variants":["Visual glyphs build manuscript stemmas","Glyph shapes alone build manuscript stemmas","Visual shapes reconstruct manuscript stemmas","Image glyphs form manuscript stemmas"]},"model":"grok-4.3","cost_usd":0.007351,"raw_usage":{"total_tokens":3443,"prompt_tokens":790,"num_sources_used":0,"completion_tokens":46,"cost_in_usd_ticks":73512000,"prompt_tokens_details":{"text_tokens":790,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2607,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":790,"tokens_out":46,"duration_ms":37899,"temperature":1.0,"reasoning_tokens":2607,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T16:01:44.604070+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the visual pipeline on a corpus whose copying relationships are already known from textual collation and finding that the resulting stemma differs substantially from the established tree.","supporting_citations":[],"review_version":1}