Pith. sign in

REVIEW 4 major objections 6 minor 24 references

When outcomes arrive late, deterministic code can own a provisional ranking of model-proposed strategic routes and be graded only after the world resolves.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 07:56 UTC pith:IO2LFTFI

load-bearing objection Carefully scoped protocol pilot for code-owned ranking under delayed ground truth; honest null ablation and leakage contrast, but tiny reconstructed n keeps the result feasibility-only. the 4 major comments →

arxiv 2607.10972 v1 pith:IO2LFTFI submitted 2026-07-13 cs.AI

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

classification cs.AI
keywords code-owned evaluationdelayed ground truthtyped strategic routesprovisional forecast-rankingproposal-authority separationidentity leakageretrospective pilotventure route selection
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most model evaluations either check a contract available at scoring time or use feedback that arrives inside the loop. This paper studies the harder case where ground truth is delayed, censored, or private, so code cannot verify correctness yet and must instead issue an auditable provisional forecast. RouteCast does this for competing typed strategic routes: models propose routes and factors; point-in-time evidence, frozen priors, and versioned deterministic transforms produce a ranking and a binding next test; later outcomes score the forecast. A small retrospective venture pilot on 21 binary cases shows preliminary whole-packet discrimination and a clear identity-exposure leakage warning for LLM judges, while a preregistered typed-route decomposition adds no ranking edge. The point is feasibility and integrity of the authority boundary, not a claim of calibrated probabilities or better decisions yet.

Core claim

Under delayed, censored, or private ground truth, deterministic code can own a provisional forecast-ranking of competing model-generated typed strategic routes from point-in-time evidence and frozen transformations. In a 21-case binary retrospective venture pilot the whole-packet RouteCast score showed preliminary discrimination (AUC 0.756), identity exposure raised apparent LLM-judge performance, and a preregistered typed-route decomposition was indistinguishable from the whole-packet score and a simple heuristic.

What carries the argument

RouteCast: an authority-boundary protocol in which models only propose routes and factors; admissible point-in-time evidence, frozen priors, and versioned deterministic arithmetic produce the provisional forecast-ranking (staged expected value, evidence reweighting, risk gates, binding transition) that later outcomes evaluate.

Load-bearing premise

That reconstructed, name-masked historical company packets still capture the real decision-time information set without hindsight contamination or residual recognizability.

What would settle it

A preregistered prospective cohort with frozen typed transitions and milestone predicates before any outcomes: if the code-owned ranking fails to beat strong structured baselines on both discrimination and calibration (Brier, ECE) once outcomes resolve, the provisional-forecast claim does not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper formalizes a regime in which deterministic code must issue a code-owned provisional forecast-ranking when ground truth is delayed, censored, or private, and instantiates it as RouteCast for competing model-generated typed strategic routes. Models propose routes and factors; point-in-time evidence, frozen priors, and versioned deterministic transformations produce a ranking that is later scored against outcomes. A retrospective venture pilot on 21 binary YC 2012–2014 cases reports whole-packet AUC 0.756 [0.471, 0.980], a blind LLM judge at 0.678, an identity-exposed judge at 0.761, and a preregistered typed-route decomposition ablation that is indistinguishable from the whole-packet score and a deterministic heuristic. Claims are carefully scoped as an auditable feasibility and integrity result, not prospective calibration, decision utility, decomposition advantage, or cross-domain validity.

Significance. If the regime formulation and protocol hold, the paper supplies a useful missing category between checkable contracts (GroundEval/CARE-style) and model-owned delayed scoring (DIALECTIC, DeLLMa, venture LLM judges): an explicit authority boundary in which models may propose and extract but versioned code owns the forecast-ranking under delayed ground truth. Strengths include honest null ablation reporting, frozen score artifacts with SHA-256 prefixes and preregistered tie-or-loss rules, an identity-exposure leakage contrast, and diagnostics that surface failure modes (TAM dominance, E0 inflation, flat product anti-ranking, adversarial omission). These are genuine protocol-engineering contributions even when the empirical discrimination claim is weak. The work is best read as a methods/protocol paper with a small integrity pilot rather than as a forecasting-performance result.

major comments (4)
  1. §6.1–6.2 and Limitations (Internal Validity): The strongest empirical claim—whole-packet retrospective discrimination (Table 2, AUC 0.756)—depends on reconstructed point-in-time packets approximating the admissible decision-time information set after masking. The paper itself notes hindsight contamination risk, residual recognizability, and that the cohort is not cleanly held out from prior development; Appendix Table A1 flags Recog. risk = yes on several cases including positives, and the identity-exposed contrast (0.678 → 0.761) shows leakage sensitivity on the same material. For the AUC to support even “preliminary” evidence of the delayed-ground-truth regime, the manuscript needs either a stronger leakage audit (e.g., independent rater recognition rates, source-date provenance checks, sensitivity to dropping high-recog cases) or an explicit demotion of discrimination to a secondary i
  2. Table 2 / §7.1: With n = 21 and only 6 positives, the company-bootstrap CI [0.471, 0.980] includes chance-level performance. Calling this “preliminary retrospective discrimination” is load-bearing for the feasibility narrative but is statistically consistent with no discrimination. The text should either (a) reframe the pilot as an integrity/process demonstration without a discrimination claim, or (b) report a pre-specified minimum effect size / decision rule under which the pilot would be judged non-informative, and apply it. Wide CIs alone do not make a positive discrimination claim safe.
  3. §5.3 and §8 (decomposition diagnostics): Staged-NPV discrimination was dominated by the terminal value proxy V = TAM × α_capture × α_defensibility rather than by the transition probability chain; the isolated chain product anti-ranked (AUC 0.333). This is a construct-validity problem for the central evaluation object—“competing typed strategic routes”—because the ranking may largely recover a large-market signal available without route structure. The manuscript should quantify how much of whole-packet and decomposed ranking is explained by TAM/terminal value alone (e.g., partial correlations or a TAM-only baseline in Table 2/3) and state whether residual route-structure signal remains after controlling for it.
  4. §5.1–5.2 free parameters: Evidence reweighting (W_reality clips [0.3, 1.5], 0.5 coefficients, saturation form), criterion weights w_i, and α_capture/α_defensibility are versioned but not sensitivity-analyzed on the pilot. Because the pilot is offered as integrity evidence for a code-owned pipeline, at least a one-at-a-time or leave-one-component-out sensitivity on the frozen 21-case subset is needed to show that the reported AUC is not an artifact of a particular prior/weight choice. Without that, “code-owned” is auditable in form but not shown to be robust in content.
minor comments (6)
  1. Table 1: “GroundEval / CARE” are grouped in one row despite different outcome-resolution regimes (checkable-now vs in-loop). Splitting them would sharpen the residual-relationship claim in §3.5.
  2. §5.4: The binding-transition heuristic b = arg max_i U_i D_i is introduced without a formal definition of D_i (downstream stake). A one-line definition or pointer would help reproducibility.
  3. Figure 1 vs §5: Notation for learning cost is ℓ_b in the figure caption and ℓ_i in the text; unify.
  4. §7.2: The paper correctly notes that whole-packet RouteCast and the identity-exposed judge are not established as statistically different, but still juxtaposes 0.756 vs 0.761 in the abstract. Soften the abstract juxtaposition or add the missing paired contrast as a limitation callout.
  5. Appendix Table A1: “Recog. risk” is binary and sparsely annotated; a short methods note on how recognition risk was assigned would aid interpretation of the leakage discussion.
  6. References include several 2026 arXiv preprints; ensure citation versions and titles match the public records at camera-ready time.

Circularity Check

0 steps flagged

No significant circularity; provisional rankings are frozen code outputs from point-in-time inputs, later scored against separately stored outcomes.

full rationale

The paper formalizes a protocol (Sections 2, 4–5) in which models propose typed routes/factors while versioned code owns the forecast-ranking from admissible It0 evidence, frozen priors/reference classes, and deterministic transformations (e.g., evidence reweighting, staged EV, binding-transition heuristic). The retrospective pilot freezes whole-packet and decomposition scores before outcome join (Section 6.3: scores frozen 2026-07-07 / commit b0dbd8a with SHA-256 prefixes; preregistration at cb7625d), so reported AUCs (Table 2, Section 7) and the null ablation (Table 3, Section 8) are post-hoc evaluations of pre-specified rankings, not quantities forced by construction from the binary labels. Priors are taken from a frozen reference-class table rather than fitted to pilot outcomes (Section 5.2); terminal-value proxies are explicitly uncalibrated ranking devices (Section 5.3). No load-bearing self-citations, uniqueness theorems imported from the author, ansatzes smuggled via prior author work, or renaming of known results appear. Residual hindsight/recognizability risks in packet reconstruction are internal-validity/leakage concerns (Limitations, Appendix A1), not definitional circularity of the derivation chain. The pilot is scoped only as an integrity/feasibility audit.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 3 invented entities

The central feasibility claim rests on standard decision-analysis scaffolding plus several protocol-specific proxies and reconstruction assumptions. Free parameters are the hand-chosen reweighting clips/coefficients, terminal-value factors, and binding-transition heuristic, not fitted to the 21 outcomes in the reported analysis. Invented entities are protocol constructs (typed route packet, code-owned forecast-ranking, binding transition) with independent handles only via later outcome resolution, which the pilot only partially exercises.

free parameters (4)
  • Evidence reweighting coefficients and clips (0.5, sat form, W_reality in [0.3,1.5], p clip [0,1])
    Section 5.2 defines a saturating multiplicative adjustment of frozen priors; coefficients and bounds are protocol choices, not derived from the pilot outcomes.
  • Terminal value proxy factors alpha_capture and alpha_defensibility in V = TAM × α_capture × α_defensibility
    Section 5.3; frozen ranking proxy whose dominance is later reported as a construct-validity limitation in the decomposition ablation.
  • Binding-transition selector b = arg max_i U_i D_i
    Section 5.4 VoI-inspired heuristic; not a calibrated knowledge-gradient or cost-aware VoI rule.
  • Criterion weights w_i and clip-to-[0,100] sub-scores in aggregate T
    Section 5.1; aggregate arithmetic is code-owned but weights/criteria are versioned design choices.
axioms (5)
  • domain assumption Expected-utility / staged real-options folding of route value with abandonment-aware later costs
    Sections 3.5 and 5.3 cite Savage/Raiffa–Schlaifer and Dixit–Pindyck; used as the ranking objective, not re-derived.
  • domain assumption Reference-class / outside-view priors can be frozen and used as p_prior(e) without accepting model prose priors
    Sections 3.5 and 5.2; Kahneman–Lovallo outside view invoked as governance principle.
  • ad hoc to paper Point-in-time reconstructed public packets for YC 2012–2014 approximate admissible decision-time information after masking
    Section 6; load-bearing for interpreting retrospective AUC as integrity evidence.
  • ad hoc to paper Binary later-outcome mapping of heterogeneous venture trajectories is a valid discrimination target for route quality
    Sections 6.1 and 10 Construct Validity; funding/traction/persistence compressed into positive/negative labels.
  • domain assumption Proposal–authority separation remains meaningful when the authority issues a forecast rather than a present-time safety check
    Sections 3.1 and 12.1; extends Simplex/runtime-assurance pattern into delayed-truth regime.
invented entities (3)
  • Typed strategic route / transition object e_i with p_i, U_i, costs, kill conditions, binding transition no independent evidence
    purpose: Provide a structured evaluation object for competing model-generated plans and later transition-level resolution
    Defined in Sections 2 and 5; protocol construct rather than a physical entity; independent evidence would be prospective transition resolution, not yet established.
  • Code-owned provisional forecast-ranking under delayed ground truth no independent evidence
    purpose: Name the authority boundary when deterministic code cannot check correctness at scoring time
    Core regime formulation in Abstract/Introduction; residual combination of known pieces; falsifiable via later calibration and utility studies.
  • Packetized provenance (proposal / evidence / ranking packets) independent evidence
    purpose: Make leakage checks, freezes, and later resolution auditable
    Section 4 operational structure; engineering construct with audit utility independent of discrimination performance.

pith-pipeline@v1.1.0-grok45 · 16221 in / 3659 out tokens · 26596 ms · 2026-07-14T07:56:13.950608+00:00 · methodology

0 comments
read the original abstract

Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast. RouteCast instantiates this regime for model-generated typed strategic routes: models propose candidate routes and structured factors; point-in-time evidence, reference classes, and deterministic transformations produce a provisional forecast-ranking; later outcomes evaluate the forecast. In a retrospective venture pilot on 21 binary-outcome cases (6 positive, 15 negative), the whole-packet RouteCast score showed preliminary retrospective discrimination (AUC 0.756, 95% CI [0.471,0.980]), while a blind LLM judge reached AUC 0.678 [0.419,0.897] and an identity-exposed LLM judge reached AUC 0.761 [0.515,0.944], consistent with recognition- or outcome-related leakage risk. A preregistered decomposition ablation on the same binary subset found that converting the identical inputs into typed staged routes was indistinguishable from the whole-packet score (Delta AUC = -0.144, 95% CI [-0.471,0.176]) and from a deterministic heuristic (Delta AUC = -0.089, 95% CI [-0.412,0.278]). The pilot establishes an auditable feasibility result and exposes failure modes; it does not establish prospective calibration, causal decision improvement, route-decomposition advantage, or cross-domain validity.

Figures

Figures reproduced from arXiv: 2607.10972 by Aleh Manchuliantsau.

Figure 1
Figure 1. Figure 1: Typed strategic route. Each transition ei carries a provisional success probability pi , epistemic uncertainty Ui , and an execution cost c E i . A selected binding transition can be tested at learning cost ℓi before full execution. The terminal node carries value V . A strategic route is a candidate path from a current state to a target state. A typed transition is one edge in that path with a source stat… view at source ↗
Figure 2
Figure 2. Figure 2: RouteCast authority-boundary architecture. Models propose routes and factors; point-in-time evidence, [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

24 extracted references · 6 linked inside Pith

  1. [1]

    DIALECTIC: A multi- agent system for startup evaluation

    Jae Yoon Bae, Simon Malberg, Joyce Galang, Andre Retterath, and Georg Groh. DIALECTIC: A multi- agent system for startup evaluation. InProceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), pages 711–727, Rabat, Morocco, March 2026. Association for Computational Linguis- tics...

  2. [2]

    Look-ahead-bench: a standard- ized benchmark of look-ahead bias in point-in-time LLMs for finance, 2026

    Mostapha Benhenda. Look-ahead-bench: a standard- ized benchmark of look-ahead bias in point-in-time LLMs for finance, 2026. arXiv:2601.13770 [cs.AI]

  3. [3]

    Glenn W. Brier. Verification of forecasts expressed in terms of probability.Monthly Weather Review, 78(1):1–3, 1950. https://journals.ametsoc.org/ view/journals/mwre/78/1/1520-0493_1950_078_ 0001_vofeit_2_0_co_2.xml

  4. [4]

    Value- BlindBench: Agreement-gated stress testing of LLM- judged investment rationales before returns are ob- servable, 2026

    Sidi Chang, Peiying Zhu, and Yuxiao Chen. Value- BlindBench: Agreement-gated stress testing of LLM- judged investment rationales before returns are ob- servable, 2026. arXiv:2604.25224 [cs.AI]

  5. [5]

    VCBench: Benchmarking LLMs in venture capital,

    Rick Chen, Joseph Ternasky, Afriyie Samuel Kwesi, Ben Griffin, Aaron Ontoyin Yin, Zakari Salifu, Kelvin Amoaba, Xianling Mu, Fuat Alican, and Yigit Ihlamur. VCBench: Benchmarking LLMs in venture capital,

  6. [6]

    arXiv:2509.14448 [cs.AI]

  7. [7]

    Robert G. Cooper. Stage-gate systems: A new tool for managing new products.Business Horizons, 33(3):44– 54, 1990

  8. [8]

    Csaszar, Aticus Peterson, and Daniel Wilde

    Felipe A. Csaszar, Aticus Peterson, and Daniel Wilde. The strategic foresight of LLMs: Evidence from a fully prospective venture tournament, 2026. arXiv:2602.01684 [econ.GN]

  9. [9]

    Dixit and Robert S

    Avinash K. Dixit and Robert S. Pindyck.Invest- ment under Uncertainty. Princeton University Press, Princeton, NJ, 1994

  10. [10]

    GroundEval: A deterministic replace- ment for LLM-as-Judge in stateful agent evaluation,

    Jeffrey Flynt. GroundEval: A deterministic replace- ment for LLM-as-Judge in stateful agent evaluation,

  11. [11]

    arXiv:2606.22737 [cs.AI]

  12. [12]

    Justin G. Fuller. Run-time assurance: A rising tech- nology. In2020 IEEE/AIAA 39th Digital Avionics Systems Conference (DASC), pages 1–9. IEEE, 2020

  13. [13]

    Tilmann Gneiting and Adrian E. Raftery. Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477):359–378, 2007

  14. [14]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural net- works. InProceedings of the 34th International Con- ference on Machine Learning, volume 70 ofProceed- ings of Machine Learning Research, pages 1321–1330. PMLR, 2017. https://proceedings.mlr.press/ v70/guo17a.html

  15. [15]

    Ronald A. Howard. Information value theory.IEEE Transactions on Systems Science and Cybernetics, 2(1):22–26, 1966

  16. [16]

    Timid choices and bold forecasts: A cognitive perspective on risk taking.Management Science, 39(1):17–31, 1993

    Daniel Kahneman and Dan Lovallo. Timid choices and bold forecasts: A cognitive perspective on risk taking.Management Science, 39(1):17–31, 1993

  17. [17]

    CARE: Controlling LLM-generated policies through auditable review of evidence in scientific experimentation, 2026

    Guanyu Liu, Weiyi Kong, Zeyu Wang, Boer Zhang, Baiqing Li, Peiyu Zhang, and Tianyu Shi. CARE: Controlling LLM-generated policies through auditable review of evidence in scientific experimentation, 2026. arXiv:2606.14581 [cs.LG]

  18. [18]

    DeLLMa: Decision making un- der uncertainty with large language models, 2024

    Ollie Liu, Deqing Fu, Dani Yogatama, and Willie Neiswanger. DeLLMa: Decision making un- der uncertainty with large language models, 2024. arXiv:2402.02392 [cs.AI]

  19. [19]

    Harvard Business School, Boston, 1961

    Howard Raiffa and Robert Schlaifer.Applied Statisti- cal Decision Theory. Harvard Business School, Boston, 1961

  20. [20]

    Savage.The Foundations of Statistics

    Leonard J. Savage.The Foundations of Statistics. John Wiley & Sons, New York, 1954

  21. [21]

    D. Seto, B. Krogh, L. Sha, and A. Chutinan. The simplex architecture for safe online control system up- grades. InProceedings of the 1998 American Control Conference, volume 6, pages 3504–3508. IEEE, 1998

  22. [22]

    Using simplicity to control complexity.IEEE Software, 18(4):20–28, 2001

    Lui Sha. Using simplicity to control complexity.IEEE Software, 18(4):20–28, 2001. 8

  23. [23]

    SSFF: Investigating LLM predictive capabilities for startup success through a multi-agent framework with enhanced explainability and performance, 2024

    Xisen Wang, Yigit Ihlamur, and Fuat Alican. SSFF: Investigating LLM predictive capabilities for startup success through a multi-agent framework with enhanced explainability and performance, 2024. arXiv:2405.19456 [cs.AI]

  24. [24]

    Separating diagnosis from control: Auditable pol- icy adaptation in agent-based simulations with LLM- based diagnostics, 2026

    Shaoxin Zhong, Yuchen Su, and Michael Witbrock. Separating diagnosis from control: Auditable pol- icy adaptation in agent-based simulations with LLM- based diagnostics, 2026. arXiv:2603.22904 [cs.AI]. 9 A Supplementary Case-Level Pilot Table The full case-level pilot table is appendix-only and uses blinded IDs. Original company names and the blind-ID map ...