{"id":"93614cd4-da2f-4150-8ec8-9a02db8798c5","arxiv_id":"2507.17212","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper claims that Goodman's grue can be ruled out because a direct measurement must give one convex error region, making grue-like formulations more complex than simple ones.","lead":"A philosophy of science paper argues that Goodman's new riddle of induction can be solved by combining the requirement that direct measurements have convex error regions with a formulation-independent measure of model complexity. The proposed solution turns the riddle into a concrete model selection rule, which the author applies to grue, conspiracy theories, and AI-era model choices.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The green/grue asymmetry is produced by Postulate 1's convexity and Definition 3's assumed non-triviality; both are stipulated rather than proved, so the solution may be definitional rather than discovered.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing spot: Postulate 1 and Definition 3 are both asserted rather than derived, and the grue asymmetry depends on them. I agree with the CONDITIONAL verdict because the paper is transparent about these postulates, offers a testable criterion (counterexamples from scientific consensus), and does not hide the definitional character of Postulate 1. My concern does not move the verdict: the proposal is conditional on the postulates being independently justified, which the paper has not yet supplied. The concrete test I propose would settle whether Definition 3 is non-trivial in a finite, controlled setting; if it is, the remaining work is to justify Postulate 1 empirically. I am not raising an objection about the author's self-citations or the philosophical framing; those are not the load-bearing weakness. The core issue is that the central solution's two pillars—convex direct measurements and non-trivial epistemic complexity—are stipulated, not established, so the solution could be read as re-describing Goodman's riddle rather than solving it.","tokens_in":17065,"tokens_out":11290,"duration_ms":137858,"concrete_test":"Formalize the framework in a finite first-order language with a fixed syntax for measurement axioms and a fixed symbol-count length measure, then enumerate (up to a bounded size) all formulations satisfying Definition 2's logical and empirical equivalence for a toy Goodman case (e.g., 'all emeralds are green' vs. 'all emeralds are grue' with a finite observation record). Compute Definition 3's epistemic complexity for both models, including formulations that introduce new directly measurable primitives with convex error boxes. If any empirically equivalent grue formulation has length no greater than the shortest green formulation, the claimed asymmetry fails; if no such formulation exists, the non-triviality of C and the complexity gap are demonstrated in this concrete fragment, isolating where an external justification of Postulate 1 would still be needed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that convexity plus epistemic complexity rules out grue rests on two unproved assumptions. First, Postulate 1 (Sec. 2.1) declares that a valid single direct measurement always yields a convex error region; the split region in Fig. 2/3 is then inadmissible by fiat. The paper explicitly calls Postulate 1 an 'implicit (partial) definition of direct measurements,' which makes the green/grue asymmetry partly analytic: grue is not directly measurable because direct measurability is defined to exclude split regions. Since the paper's own adequacy test (Sec. 3) is descriptive accuracy against scientific consensus, the postulate must be empirically checkable, yet no independent characterization of 'direct' is given, so a proponent of grue can stipulate a grue-meter whose single-readout distribution is bimodal and insist it is direct. Second, Definition 3 (Sec. 2.2) defines epistemic complexity as the minimum length over all logically and empirically equivalent formulations and claims this minimum is 'in general, not trivial anymore'; no proof is supplied. The min over arbitrary languages can be ill-defined, and it is not obvious that a carefully chosen directly measurable quantity (with a convex error box by construction) could not encode the model's full empirical content in a single short axiom, collapsing C to a constant. Both assumptions are load-bearing: if either fails, the complexity gap between green and grue disappears and the solution does not go through.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a solution to Goodman's new riddle of induction by combining two ideas: (i) a postulate that the outcome of a single direct measurement is always a central value with a convex error region, and (ii) an \"epistemic complexity\" defined as the minimum length of the assumptions over all logically and empirically equivalent reformulations. The author argues that a grue-like formulation cannot be empirically equivalent to the standard green formulation because measuring grue requires a non-convex, split error-box, so any grue reformulation is more complex without empirical advantage. The paper also argues that this combination solves the riddle in the sense of identifying the hidden assumptions behind scientific model selection, and it offers historical and conspiracy-theory examples as illustrations.","tokens_in":17317,"tokens_out":6522,"duration_ms":71813,"significance":"If the central claims hold, the paper would provide a principled, reformulation-invariant criterion that rules out grue-like predicates without appealing to psychological entrenchment or subjective simplicity, and it would connect a classic philosophical riddle to practical questions in AI model selection. The paper is clearly written, engages seriously with the literature on conceptual spaces and complexity, and explicitly identifies the adequacy conditions for a solution. The main contribution is conditional, however: the two load-bearing assumptions—convexity as a constitutive feature of direct measurements and the non-triviality of the minimum in Definition 3—are asserted rather than demonstrated. The paper's descriptive examples are suggestive but do not yet provide independent evidence for those assumptions.","major_comments":[{"comment":"The central asymmetry between green and grue rests on excluding split error regions from the outcome of a single direct measurement. The author explicitly acknowledges that Post. 1, together with Def. 1, offers an \"implicit (partial) definition of direct measurements\" and that \"it is up to the model to decide which are the direct measurements.\" Consequently, the claim that grue is not directly measurable is not an empirical result but a consequence of the postulate; a proponent of grue could simply stipulate that a grue-meter's single-readout distribution is bimodal and insist that it is direct. To carry the argument, the paper needs an independent characterization of directness, or at least a systematic empirical survey showing that no accepted scientific direct measurement has a non-convex error-box. The two examples given (room temperature and Fig. 2) are not enough to support the load placed on this postulate.","section":"Section 2.1, Postulate 1"},{"comment":"The epistemic complexity C(M) is defined as a minimum over \"all possible equivalent formulations (in any language)\" of length[P(M')]. The paper does not specify the length function, does not prove that the minimum exists, and does not prove that it is non-trivial for the relevant cases. The statement that restricting to logically and empirically equivalent formulations ensures that the Xi = 0 formulation is no longer legitimate, and that the shortest formulation is \"in general, not trivial anymore,\" is an assertion. Without a precise definition of the language class and a proof that no short, convex, directly measurable reformulation can encode the same empirical content, the complexity gap between green and grue could collapse, and the definition might be ill-defined or yield a constant for all models.","section":"Section 2.2, Definition 3"},{"comment":"The empirical equivalence relation used in Definition 3 requires \"same precision and same outcome\" for each measurable property. Precision is not an external fact; it is part of the model assumptions, since Definition 1 includes Delta(b) for every directly measurable quantity b in B. This creates a circularity: whether two formulations are equivalent depends on the very precision values whose effect on complexity is being assessed. The paper should specify how equivalence is to be judged independently of the model's own stipulations, or explain why this dependence does not undermine the claimed reformulation independence.","section":"Section 2.2, Definition 2"},{"comment":"The paper's justification strategy is descriptive accuracy: the model is accepted if no counterexample against scientific consensus exists. But the paper never operationalizes \"broad scientific consensus\" nor explains how to identify a counterexample independently of the framework. The Bielefeld conspiracy example is essentially a restatement of the same unavailability-of-records argument used for grue, so it does not provide independent support. The \"no counterexample\" claim is therefore a conjecture rather than a test of the model; the author should either provide a falsification protocol or soften the claim accordingly.","section":"Sections 3.2 and 5"}],"minor_comments":[{"comment":"The paper says that all important conclusions are maintained under the more general Postulate 1', but it does not demonstrate this; a short verification would be helpful, especially given that the probability-distribution formulation is what makes the convexity claim checkable.","section":"Section 2.1, Postulate 1'"},{"comment":"The term length[P(M')] is never defined. If it is intended as a string length in some formal language, that language must be specified; if it is an informal notion of amount of assumptions, the claim of precision is strained.","section":"Section 2.2, Definition 3"},{"comment":"There are several typographical errors, including \"Goodnam\" in Section 1.2, \"constrints\" in Section 2, \"explicitely\" and \"alghough\" in Section 2.1, \"Gardenfor's\" in footnote 9, and \"discipleines\" in Section 5. These should be corrected.","section":"Throughout"},{"comment":"The right panel of Fig. 3 would be more informative if the axes were labeled explicitly (e.g., wavelength and time), so that the reader can see exactly how the error-box splits in the grue/bleen representation.","section":"Figure 3"},{"comment":"The discussion of knowledge-what is interesting but seems only loosely connected to the formal definitions in Section 2; the author should state explicitly whether the irreducibility of knowledge-what is supposed to justify Postulate 1 or is merely a philosophical aside.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is original and likely to interest readers of this journal, but the core solution is only as strong as the two unproved assumptions identified above. A revision that either supplies a formal proof (or precise conditions) for the non-triviality of Definition 3 and an independent argument for Postulate 1, or explicitly reframes the paper as a conditional proposal, would make the contribution publishable. The author's claim of being the only published option for a reformulation-independent complexity should also be softened or supported with a more thorough literature check."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a clear, honest piece of philosophy of science. It argues that Gärdenfors' convexity constraint on direct measurements and the author's earlier epistemic-complexity framework (Scorzato 2013) should be combined, and that doing so rules out grue-like predicates. What's genuinely new is the explicit claim that these two ideas need each other: convexity makes complexity non-trivial by excluding the Ξ=0 trick, and complexity saves convexity from being too weak. The Bielefeld conspiracy example is a nice, accessible illustration. The paper is also refreshingly explicit about its own adequacy test: find a counterexample from scientific consensus. That gives critics a clear target.\n\nThe soft spots are real, and the stress-test note lands. Postulate 1—that a single direct measurement always yields a convex error-box—is doing more work than a postulate should. The paper itself calls it an 'implicit (partial) definition of direct measurements,' which means grue's non-directness is to some extent analytic: you've defined directness so that split error regions are inadmissible. A determined grue-proponent can stipulate a bimodal single-readout device and insist it's direct; you need an independent argument for why that's not legitimate, and the paper doesn't give one. Similarly, Definition 3's epistemic complexity assumes the minimum over all logically and empirically equivalent formulations is non-trivial, but no proof is supplied. It's not obvious a clever encoding couldn't compress the model's empirical content into one short axiom with a convex error-box, collapsing C to a constant. The 'no counterexamples' defense is also weak—absence of published counterexamples to your own prior model is not strong evidence.\n\nThat said, the paper doesn't oversell its formalism. It presents a philosophical model, not a theorem, and it explicitly invites falsification. The core combination is worth taking seriously, but the claim to have solved Goodman's new riddle is premature.\n\nThis deserves a serious referee. I'd send it to a good philosophy-of-science journal with a request that the author either prove or formally flag the non-triviality of C(M) and directly address the analyticity worry. It's the kind of paper that benefits from expert pushback, and it could become a useful reference point even if the solution doesn't ultimately hold. I'd bring it to a reading group, though I wouldn't build my own work on it yet.","headline":"A transparent synthesis of Gärdenfors and the author's own 2013 model that makes a real proposal, but the green/grue asymmetry is largely stipulated by Postulate 1, so 'solution' overstates what is demonstrated.","tokens_in":17902,"tokens_out":2184,"would_cite":false,"duration_ms":24154,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that Goodman's new riddle of induction dissolves once direct measurements are required to yield convex error-boxes and model complexity is measured by the shortest empirically equivalent formulation.","keywords":["Goodman's new riddle of induction","grue","model selection","epistemic complexity","convexity","direct measurement","underdetermination","Kolmogorov-Chaitin complexity"],"falsifier":"A single direct measurement whose reported uncertainty is a genuinely disconnected or non-convex region—a multimodal expected distribution for one measurement—would falsify Postulate 1 and collapse the complexity gap, because grue would then be as directly measurable as green. Alternatively, any case of model selection that the broad scientific community rejects but Definition 4 admits, or accepts but Definition 4 rules out, would falsify the paper's descriptive claim.","tokens_in":16783,"feed_emoji":"🟩","tokens_out":11450,"duration_ms":105498,"temperature":0.7,"pith_summary":"This paper claims that Goodman's new riddle of induction—the puzzle of why 'green' is projectible but 'grue' is not—can be solved by tying two ideas together: a direct measurement must report a central value inside a single convex error-box, and a model's complexity is the shortest formulation of its assumptions among all logically and empirically equivalent versions. Under that standard, the familiar $\\Xi=0$ trick fails: rewriting every hypothesis as one equation is not empirically equivalent to the original model, because it would require the lumped quantity $\\Xi$ to be directly measurable with a connected error-box, which grue-like quantities are not. The paper argues that this yields a simple selection rule—drop any model that is more complex and no more accurate than a competitor—that matches the model choices scientists actually make, and that this descriptive match is what a solution to the new riddle should be. The reason to care, beyond resolving a textbook paradox, is that the same trick now threatens the reliability assessment of AI models.","feed_headline":"Grue loses once direct measurements must give convex error-boxes","feed_subtitle":"A complexity measure that ignores grue-like rewrites of models reproduces the choices scientists actually make.","key_machinery":"The machinery is the pair consisting of Postulate 1 and Definition 3. Postulate 1 requires the outcome of a single direct measurement of a property $Q$ to be a central value $Q_0$ together with a convex error-box containing $Q_0$; it rules out split error regions as legitimate direct outcomes. Definition 3 defines the epistemic complexity $C(M)$ of a model $M$ as the minimum, over all logically and empirically equivalent formulations $M'\\equiv M$, of the length of the assumptions $P(M')$; here empirical equivalence (Definition 2) requires the translation to preserve the measurement outcomes and precisions of every measurable property. Together these make the $\\Xi=0$ reformulation inadmissible, because that rewriting breaks empirical equivalence and demands a directly measurable $\\Xi$ with a connected error-box, so the minimum is non-trivial and reformulation-independent. The paper's one-sentence summary of the mechanism: 'it is the constraint of convexity that enables a non-trivial notion of complexity.'","core_discovery":"The central claim is that the asymmetry between green and grue is not a matter of language or convention but of measurability: although any model can be expressed as $\\Xi=0$, no empirically equivalent formulation can make $\\Xi$ a directly measurable quantity with a central value and a connected error-box, because the error-box of a grue measurement splits at the critical time $t_0$ and, for emeralds first seen before $t_0$, remains undetermined later. The paper formalizes this through Postulate 1 (convexity of single direct measurement outcomes) and Definition 3 (epistemic complexity as minimum length over logically and empirically equivalent formulations), and shows that the grue model is then strictly more complex than the green model with no empirical advantage, so Definition 4 rules it out. The paper stresses that this solves the new riddle—identifying the hidden assumptions behind scientists' actual model selection—and not the old riddle of justifying induction by future success.","pith_inferences":["Extending the paper's approach: any predicate that forces split error-boxes in a shared direct-measurement basis should be empirically disfavored exactly like grue; this is a testable prediction for language design and machine-learning feature engineering.","The paper asserts rather than proves that the minimum in Definition 3 is non-trivial and attained; a formal proof, or a realistic model class where the minimum fails to be attained, would either complete or stress the foundation.","A broader research program implied here is grounding non-empirical epistemic values (simplicity, naturalness, projectibility) in measurability constraints rather than in metaphysical natural kinds or entrenchment.","For AI, the framework suggests a concrete auditing rule: a learned model that gains apparent simplicity by redefining its inputs so that measurement error-boxes split is a 'grue model' and should be downgraded; this could be operationalized as a test for shortcut learning."],"forward_implications":["The selection rule of Definition 4 gives a precise meaning to 'explaining more with less': any model that is more complex and no more accurate than a rival is ruled out, with no trade-off involved.","Because epistemic complexity is invariant under logically and empirically equivalent reformulations, simplicity comparisons remain meaningful across very different theories, including across scientific revolutions.","The same framework dismisses conspiracy theories: extra ad-hoc assumptions like '$\\Xi$-people' buy conciseness only by sacrificing empirical accuracy, because the corresponding measurements are not available.","Bayesian confirmation requires prior probabilities; the paper argues the prior choice can be anchored only by this reformulation-independent complexity measure, otherwise the $\\Xi$ trick makes priors arbitrary.","This is a solution to the new riddle, not the old one: it describes the hidden assumptions behind scientists' actual choices rather than promising to justify the future success of science."],"supporting_citations":[{"why":"Poses the new riddle of induction that the paper aims to solve.","marker":"(Goodman, 1955)"},{"why":"Supplies the core idea that natural properties form convex sets in conceptual spaces, which the paper narrows to directly measurable properties.","marker":"(Gärdenfors, 1990)"},{"why":"Introduces the combination of simplicity with measurability constraints that the paper extends into epistemic complexity and model selection.","marker":"(Scorzato, 2013)"},{"why":"Provides Kolmogorov-Chaitin complexity, the inspiration for Definition 3, which the paper modifies by restricting to empirically equivalent formulations.","marker":"(Kolmogorov, 1965)"},{"why":"Formalizes algorithmic randomness and complexity that underlies the KC measure the paper adapts.","marker":"(Chaitin, 1975)"},{"why":"Relates knowledge-what and induction; the paper positions its own convexity-based approach against this account.","marker":"(Gärdenfors and Stephens, 2017)"}],"fun_headline_variants":["Grue fails when complexity includes direct measurement error-bounds","New solution to Goodman's riddle: grue is more complex, not just odd","Complexity plus measurability resolves Goodman's grue riddle","Why grue isn't a real alternative: error-boxes split, complexity grows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument hangs on Postulate 1, that a single direct measurement always reports a central value inside one convex error-box and never a split or disconnected region; if scientists could legitimately report a split region as a direct measurement, grue would be directly measurable and the complexity advantage of green disappears.","fun_headline_variants_meta":{"raw":{"variants":["Grue fails when complexity includes direct measurement error-bounds","New solution to Goodman's riddle: grue is more complex, not just odd","Complexity plus measurability resolves Goodman's grue riddle","Why grue isn't a real alternative: error-boxes split, complexity grows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2168,"prompt_tokens":911,"completion_tokens":1257,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":1177}},"tokens_in":527,"tokens_out":1257,"duration_ms":8932,"temperature":1.0,"reasoning_tokens":1177,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:54:51.431636+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single direct measurement whose reported uncertainty is a genuinely disconnected or non-convex region—a multimodal expected distribution for one measurement—would falsify Postulate 1 and collapse the complexity gap, because grue would then be as directly measurable as green. Alternatively, any case of model selection that the broad scientific community rejects but Definition 4 admits, or accepts but Definition 4 rules out, would falsify the paper's descriptive claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Poses the new riddle of induction that the paper aims to solve."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the combination of simplicity with measurability constraints that the paper extends into epistemic complexity and model selection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Kolmogorov-Chaitin complexity, the inspiration for Definition 3, which the paper modifies by restricting to empirically equivalent formulations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes algorithmic randomness and complexity that underlies the KC measure the paper adapts."}],"review_version":1}