{"id":"9fc618b1-4941-4940-bd56-e2549936f148","arxiv_id":"2509.25359","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper's abstract claims that Schatten Norm and MOM reflect output length and that geometric features add modest classifier accuracy over text statistics, but the body instead reports consistent generator rankings from intrinsic dimensionality and effective rank without the promised length…","lead":"This preprint tests whether geometric properties of LLM internal representations, such as intrinsic dimensionality and effective rank, can rank the quality of generated text. The abstract promises a controlled stress-test showing which metrics fail once text length is accounted for, but the body reports a different validation study and does not contain that analysis.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-tester ranking consistency is not shown to survive length control or to track human-judged quality, so the universal-quality claim is unsubstantiated.","rationale":"The reader's weakest assumption is that cross-tester consistency could be a shared length confound; I agree this is a serious threat, and it is strengthened by the submitted abstract's own finding that Schatten Norm and MOM collapse under length control. I would not fully agree that it is the only load-bearing point: even if length is controlled, the paper never validates the resulting ranking against human judgments of quality, and Figure 4's correlations are computed over 9 aggregated points with no p-values (the paper itself says more observations would be required). The submitted version is internally inconsistent: the abstract describes a length-controlled stress-test and a classifier comparison that do not appear in the body, while the body's abstract and Section 4.1 make a stronger universal-quality claim. Because the submitted abstract's own stress-test finding undermines the body's central inference for some metrics and no human-level validation is reported, the central claim is not established. The reader's REJECT verdict is appropriate; a revision adding length-matched results, confidence intervals, human ratings, and code and data could move to CONDITIONAL.","tokens_in":19083,"tokens_out":6447,"duration_ms":58541,"concrete_test":"Recompute the Section 4.1 cross-tester Spearman matrix (Figure 16) after truncating or padding every text to a fixed token count, and again after partialling out log token length; if the minimum inter-tester rho drops below 0.9, or the recommended metrics' generator ranking changes materially, the consistent-ranking evidence is length-driven rather than quality-driven. Report the same matrix with permutation-based confidence intervals over the 8 to 10 generator points.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract; Section 4.1) is that consistent rankings across tester models show that ERank, MEV, and CorrInt are universal reference-free quality proxies. This requires the shared ranking not to be an artifact of a common confound, and the rank order to correspond to quality. The first condition is directly threatened by the submitted abstract's own finding that Schatten Norm and MOM mainly reflect output length and collapse under length control, and by Table 3, where average generator lengths vary from 16.18 to 22.22 tokens despite the prompt's length-matching instruction. Every metric in Section 3.5 is computed on a token-representation matrix whose row count equals sequence length, so singular-value and correlation-dimension estimates are structurally length-sensitive; no length-matched or partial-correlation analysis is reported for the recommended metrics. The second condition is also unvalidated: Figure 4 correlations use 9 aggregated rows with no p-values, no human preference judgments are reported, and the submitted abstract's classifier result (78% versus 69%) targets generator identification, not quality. Section 5 concedes that absolute reliability requires further validation. The inference from testers agree to metrics measure text quality is therefore load-bearing and unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies eight geometric metrics derived from LLM internal representations (e.g., Maximum Explainable Variance, Effective Rank, Intrinsic Dimensionality, Schatten norms, CorrInt) as candidate reference-free quality measures. Using six tester models and eight generator models on a paraphrase/rewrite task in English, German, and Russian, it reports that tester models produce consistent rankings of generators, that some geometric metrics correlate with established text-quality metrics (BLEURT, MAUVE, GPT perplexity, compression ratio), and it recommends ERank, MEV, and CorrInt as efficient reference-free quality proxies. The full-text abstract additionally claims that Intrinsic Dimensionality and Effective Rank are universal assessments of text naturalness and quality, while the official arXiv abstract promises findings on length effects and a generator-classification experiment, neither of which appears in the body.","tokens_in":19233,"tokens_out":4784,"duration_ms":40867,"significance":"If the universal-quality claim were supported, the contribution would be practically important: a reference-free, annotation-free evaluation of generated text using small tester models. The paper has notable strengths: it evaluates a diverse set of tester models including a diffusion-based LLM, covers three languages, presents precise definitions for several metrics, and reports an interesting empirical regularity of high cross-tester ranking agreement (Spearman minimum 0.947). However, the evidence falls far short of the claim. The key correlation analysis uses about nine aggregated data points without p-values; no human quality judgments are reported; the full text contains no length-matched or partial-correlation analysis for the recommended metrics; and the official abstract describes length-control and classifier results that are absent from the body. The central inference from 'testers agree' to 'metrics measure text quality' is load-bearing and unsupported.","major_comments":[{"comment":"The official arXiv abstract states that the work separates 'genuine geometric signal from text-length effects,' reports that Schatten Norm and MOM mainly reflect output length and lose discriminative power once length is controlled, and gives a classifier result (78% versus 69% accuracy on generator identification). None of these length-controlled analyses or classifier experiments appears in Sections 3–6 of the full text, and the full-text abstract instead makes the stronger claim that Intrinsic Dimensionality and Effective Rank are universal assessments of text quality. This abstract-body inconsistency is not a presentation issue: it directly affects which claims the paper is entitled to make, and the official abstract itself qualifies the universal-quality claim by showing that some geometric metrics are confounded by length.","section":"Abstract (arXiv) vs. full-text body"},{"comment":"Every geometric metric in Section 3.5 is computed on a token-representation matrix X^(l)_g of size n×d, where n is the sequence length, so singular-value spectra and correlation-dimension estimates are structurally sensitive to n. Table 3 shows that despite the length-matching prompt of Section 3.2, average output lengths range from 16.18 to 22.22 tokens, with Deepseek-R1 showing a standard deviation of 11.51 tokens. No length-matched comparison or partial-correlation analysis controlling for length is reported for the recommended metrics (ERank, MEV, CorrInt). Under these conditions, the consistent cross-tester ranking of generators reported in Section 4.1 may reflect shared length differences rather than 'inherent text characteristics,' and the central claim is therefore threatened by a confound that the paper itself identifies in the official abstract but does not address in the body.","section":"Section 3.4 and Table 3"},{"comment":"The only direct evidence linking geometric metrics to text quality is the Spearman correlation matrix in Figure 4, computed over aggregated scores for eight generator models plus the original text, i.e., nine points. The text itself states that p-values are not reported because 'more observations would be required.' With n=9, correlations such as ρ=0.81 (MEV vs. GPT-PPL) and ρ=−0.76 (ERank vs. GPT-PPL) have very wide confidence intervals, and no measure of uncertainty is given. Section 4.4 nevertheless concludes that ERank, MEV, and CorrInt are 'efficient, reference-free proxies for generation quality.' This conclusion is load-bearing for the paper's universal-quality claim and is not supported by the statistical evidence presented. The paper's own limitation section (Section 5) concedes that absolute reliability requires further validation.","section":"Section 4.4 and Figure 4"},{"comment":"The recommended metric subset (ERank, MEV, CorrInt) is selected from the same correlation and ranking tables that are then used to justify the recommendation, without any hold-out evaluation, cross-validation, or external validation. Since the selection criterion and the supporting evidence are the same data, the claim that these particular metrics 'work' as quality proxies is an in-sample selection result. This does not make the paper circular in the sense of fitting parameters to labels, but it means the recommendation is not tested against data not used in its formation.","section":"Section 4.4 and Table 1"}],"minor_comments":[{"comment":"The model names are inconsistent: Section 3.1 lists 'Gemma-1-7b,' while Section 4.1 and Figure 16 use 'Gemma-1-2B'; please reconcile the naming.","section":"Section 3.1"},{"comment":"The metrics MOM and MADA are used in the tables and figures but are never defined in Section 3.5 or Table 2; please provide definitions or explicit references.","section":"Section 3.5"},{"comment":"CorrInt is named in Table 2 and used in the analysis but no formula or estimation description is given in the main text; a precise definition is needed for reproducibility.","section":"Section 3.5"},{"comment":"The 'Average' column contains a tie (Mistral-7b-it and Gemma-2b-it both have 3.0), but the rule for aggregating the metric-specific ranks is not described; please specify the averaging procedure.","section":"Table 1"},{"comment":"The caption for Figure 3 says 'This Spearman correlation demonstrates similarity among different geometric R scores,' but the main text refers to this figure as the pairwise correlation matrix; the wording should be aligned with the figure's actual content.","section":"Figure 3 caption"}],"recommendation":"reject","confidential_remarks":"The discrepancy between the official arXiv abstract and the full-text body is serious and should be investigated: the abstract reports length-control and classifier results that are entirely absent from the paper, while the body's abstract makes a stronger universal-quality claim than the results support. This suggests the submitted PDF and the abstract may be from different versions of the work. Beyond that, the central scientific claim is not supported by the evidence: the quality-correlation analysis uses roughly nine aggregated points without p-values, there is no length-controlled analysis in the body, and there are no human quality judgments. These are not local presentation fixes; they require new experiments and a substantial revision of the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2509.25359. First, the abstract on the arXiv listing and the full text are different papers. The listing abstract describes a length-controlled stress-test, a classifier comparison (78% vs 69% for generator identification), and concludes that geometric metrics show only moderate association with lexical diversity and that failure detection is the most promising near-term use. The body claims that Intrinsic Dimensionality and Effective Rank are universal assessments of text naturalness and quality, and that cross-tester rank consistency proves the metrics capture intrinsic text properties. The body never reports the length-collapse or classifier analyses.\n\nWhat is genuinely new is the cross-tester consistency: six tester models, 0.5B to 8B, including a diffusion LLaDA, rank eight generators nearly identically on several geometric metrics (min Spearman 0.94). That is a real observation, and the tester/generator separation is a sensible design. The metric definitions are explicit, and Section 5 honestly concedes that absolute reliability needs further validation.\n\nThe soft spots are load-bearing. The abstract-body mismatch alone makes the central claim impossible to evaluate; as submitted, this is two papers. And cross-tester agreement does not establish quality. The metrics are computed on token-representation matrices whose row count equals sequence length, so they are structurally length-sensitive. Table 3 shows average generator lengths from 16.18 to 22.22 tokens despite the length-matching prompt. The listing abstract itself says Schatten Norm and MOM collapse under length control, but the body never applies length-matching or partial correlation to the recommended ERank/MEV/CorrInt. The shared rankings could be a shared length artifact.\n\nThe correlations with quality metrics in Figure 4 also rest on nine aggregated points, no p-values, and no human judgments; the authors say as much. Recommending ERank, MEV, and CorrInt from the same table that validates them adds a selection problem. These problems don't kill the underlying idea, but they do undercut the submission as written.\n\nThis paper is worth engaging with, not citing. The cross-tester consistency deserves an explanation, and a length-controlled version would be a real contribution. I would send it to peer review with a strong request for major revision: reconcile the two versions, report the length-controlled results, and provide bootstrap intervals or p-values. I would not cite it in its current form.","headline":"The body and the listing abstract are effectively two different papers, and the cross-tester consistency result—real but length-sensitive—does not support the universal-quality claim.","tokens_in":19857,"tokens_out":5336,"would_cite":false,"duration_ms":46790,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Geometric fingerprints of LLM internal states can rank text quality with no reference text needed.","keywords":["intrinsic dimensionality","effective rank","text quality evaluation","reference-free evaluation","LLM internal representations","anisotropy","text naturalness","geometric metrics"],"falsifier":"Generate rewrites of the same reviews with token counts matched across all eight generators, recompute layer-averaged ERank, MEV, and CorrInt, and check whether human-vs-synthetic separation and the generator ranking survive; if the rankings collapse or reverse, the signal is length, not naturalness.","tokens_in":18824,"feed_emoji":"📐","tokens_out":8134,"duration_ms":67126,"temperature":0.7,"pith_summary":"The paper tries to establish that geometric properties of a language model's internal representations—above all Effective Rank and Intrinsic Dimensionality—can serve as reference-free proxies for how natural and high-quality generated text is. The evidence is an agreement result: six tester models, from 0.5B to 8B parameters and including one diffusion model, rank eight text generators in nearly the same order, with minimum Spearman correlation 0.947. That agreement is what lets the authors call the signal intrinsic to the text rather than an artifact of the measuring model. The paper is also a stress-test: it reports that Schatten Norm and a moment-based intrinsic-dimension estimator mainly track output length, that geometric metrics add modest information beyond text statistics (78% versus 69% accuracy for generator identification), and that the most credible near-term use is failure detection rather than absolute quality scoring.","feed_headline":"Geometry of LLM internals ranks text quality, no references needed","feed_subtitle":"Six tester models from 0.5B to 8B agree on text rankings, pointing to a text-intrinsic quality signal.","key_machinery":"The workhorse is the singular-value decomposition of the hidden-state matrix $X^{(l)} \\in \\mathbb{R}^{n \\times d}$ of a tester model's layer $l$, turned into a layer-averaged score $s^R(X_g) = (1/L)\\sum_{l=1}^{L} R(X_g^{(l)})$. The metrics doing the work are Effective Rank, the exponential of the entropy of normalized singular values; Intrinsic Dimensionality, the estimated manifold dimension of the representations; Maximum Explainable Variance, the share of variance carried by the top singular value; Resultant Length, the norm of the mean normalized token embedding; Schatten norms; and MAUVE. The identity that carries the argument is agreement: because the resulting rankings of generators are nearly identical across architecturally diverse testers, the paper concludes that the geometric signal belongs to the text, not to any single model.","core_discovery":"The central claim is that Intrinsic Dimensionality and Effective Rank, averaged over the layers of any capable tester model, measure inherent properties of the text itself, not properties of the tester. Human-written reviews show higher Effective Rank and lower Maximum Explainable Variance than LLM rewrites, and the same ordering of generators appears whether the tester is a 0.5B model or an 8B diffusion model. The paper reads this consistency as evidence that the geometric scores capture text naturalness: Effective Rank correlates negatively with GPT perplexity ($\\rho=-0.76$) and positively with BLEURT ($\\rho=0.40$), while anisotropy measures MEV and Resultant Length correlate positively with perplexity and length variability. The proposed practical conclusion is to use ERank, MEV, and CorrInt as efficient, reference-free proxies for generation quality, with a small tester model standing in for human annotation.","pith_inferences":["The agreement test should be rerun on length-matched outputs: if ERank and MEV stop separating generators once token counts are equalized, the 'text-intrinsic' interpretation collapses to a length effect.","Cross-tester Spearman correlation could be used as a cheap screening gate for any new geometric metric proposed as a quality proxy, before human evaluation is spent.","Because Effective Rank is an entropy over singular values, it may partially encode tokenization and vocabulary richness; per-token normalization could separate a genuine naturalness signal from verbosity.","The failure-detection application suggests a production pattern: monitor geometric scores of streaming generations and flag sudden shifts, an operational use the paper identifies but does not benchmark."],"forward_implications":["Text quality can be scored by running a small tester model over candidate outputs, with no reference text or human labels, using ERank, MEV, or CorrInt.","The ranking transfers across model architectures, so diffusion-based language models can serve as testers just as autoregressive models do.","Schatten Norm and MOM should be dropped as quality proxies when output length varies, since they mostly reflect length rather than naturalness.","Pairing geometric scores with ordinary text statistics improves generator identification from 69% to 78% accuracy, suggesting the geometry adds signal rather than replacing statistics.","For Russian and German, the gap between human and synthetic text is smaller than for English, so cross-lingual quality claims need further validation."],"supporting_citations":[{"why":"Gives the definition of Effective Rank as the entropy-based effective dimensionality that the paper proposes as a universal quality proxy.","marker":"Roy & Vetterli (2007)"},{"why":"Shows intrinsic dimensionality is lower for AI-generated text than human text, the pattern this work extends to internal representations.","marker":"Tulchinskii et al. (2023)"},{"why":"Supplies Maximum Explainable Variance and the anisotropy perspective used to separate synthetic from natural text.","marker":"Razzhigaev et al. (2024)"},{"why":"Defines Resultant Length, the anisotropy measure that ranks generators consistently with MEV.","marker":"Ethayarajh (2019)"},{"why":"Defines the MAUVE divergence score used as an external naturalness baseline.","marker":"Pillutla et al. (2021)"},{"why":"Provides the maximum-likelihood intrinsic dimensionality estimator among the ID metrics.","marker":"Levina & Bickel (2004)"},{"why":"Provides the IMDB movie-review data whose human originals anchor the human-vs-synthetic comparisons.","marker":"Maas et al. (2011)"}],"fun_headline_variants":["Intrinsic geometry of LLMs reads text quality sans references","ERank and MEV measure text quality, but watch length effects","Geometric metrics add 9% accuracy over text stats alone","Small tester models see text naturalness via effective rank"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For the central claim to hold, the near-identical rankings produced by different tester models must be evidence about text quality and not about a property all testers share, such as output length.","fun_headline_variants_meta":{"raw":{"variants":["Intrinsic geometry of LLMs reads text quality sans references","ERank and MEV measure text quality, but watch length effects","Geometric metrics add 9% accuracy over text stats alone","Small tester models see text naturalness via effective rank"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1230,"prompt_tokens":922,"completion_tokens":308,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":239}},"tokens_in":538,"tokens_out":308,"duration_ms":3661,"temperature":1.0,"reasoning_tokens":239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:42:42.851908+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate rewrites of the same reviews with token counts matched across all eight generators, recompute layer-averaged ERank, MEV, and CorrInt, and check whether human-vs-synthetic separation and the generator ranking survive; if the rankings collapse or reverse, the signal is length, not naturalness.","supporting_citations":[{"cited_title":"The shape of learning: Anisotropy and intrinsic dimensions in transformer- based models","cited_arxiv_id":null,"evidence_quote":"Supplies Maximum Explainable Variance and the anisotropy perspective used to separate synthetic from natural text."},{"cited_title":"How contextual are contextualized word representations? comparing the geom- etry of bert, elmo, and gpt-2 embeddings","cited_arxiv_id":null,"evidence_quote":"Defines Resultant Length, the anisotropy measure that ranks generators consistently with MEV."}],"review_version":2}