{"id":"03122ca2-eb0e-4972-8df4-f27cfada7440","arxiv_id":"2607.01153","paper_version":3,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A diagnostic framework and seed benchmark that separates four inference targets hidden inside a single safety-evaluation label, with pilot evidence that LLM judges miss safety-relevant minority outcomes.","lead":"This paper argues that AI safety evaluations should decompose pass/fail labels into separate judgments about reference, model behaviour, measurement, and taxonomy rather than compress them into one number. It ships an 18-item seed benchmark, a 54-row pilot, and LLM-judge validation showing that unvalidated judges miss the minority classes that matter.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The four-bearer decomposition is under-specified: adjudication-procedure and codebook changes are assigned to multiple bearers, so the record cannot uniquely say which inference license a change retires.","rationale":"The reader's weakest assumption concerns the empirical diagnostic power of the seed pairs, specifically that non-minimal pairs confound pair-level diagnoses of source sensitivity, scope robustness, and reference stability. That concern is real but explicitly scoped in the paper: the seed items are described as development contrasts, not strict minimal pairs; the headline statistic is renamed 'joint pair completion'; and tightening to minimal pairs is deferred to Study B. It therefore does not break the principal methodological claim. The more load-bearing problem is internal to the framework: the four-bearer typology lacks an operational rule for attributing a change to a bearer when the definitions allocate the same component to more than one bearer. Because the central claim is precisely that the typology separates inference targets and tells you which projection a change retires, an undefined boundary between reference and measurement, and between measurement and taxonomy, means the prescription cannot always be followed as stated. This is checkable without new model runs; it can be tested against the released claim register and validators with synthetic revisions. The appropriate verdict is CONDITIONAL: the contribution is valuable and largely coherent, but it should be accepted only after the bearer-attribution rules are made explicit, either by enforcing disjoint bearer definitions or by specifying a precedence/multi-bearer retirement rule.","tokens_in":33967,"tokens_out":10107,"duration_ms":101812,"concrete_test":"Apply a single versioned change to the released adjudication protocol, for example replacing the four-source disagreement-allocation rule in Section 5 with a majority-vote resolution rule, while leaving items, codebook, taxonomy, judge prompts, and expected-behaviour templates fixed. Require the claim register and validators to record which projectibility bearer(s) the change retires and what repair is required. If the schema either cannot represent 'reference retired and measurement retired' as a distinct, rule-governed outcome, or cannot specify the required repair, then the four-bearer separation is not operational for exactly the class of changes the paper says one record cannot attribute.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that evaluation results should be reported separately for reference, behavioural, measurement, and taxonomic projectibility, because 'one record can't say which bearer a given change moved' (Section 1). The paper's own bearer definitions do not support unique attribution for a natural class of revisions. Section 8 defines the reference bearer as the licensed-action relation generated by a 'frozen stipulated-regime template and output-blind adjudication procedure', while the same section defines the measurement bearer as a 'versioned scorer, codebook, implementation, adjudication procedure, criterion semantics, and enumerated prompt and evidence-access perturbations'. An adjudication-procedure change is therefore simultaneously a reference-bearer change and a measurement-bearer change, with no stated rule for which projection is retired or whether both must be re-warranted. The same overlap appears for codebooks: Section 1 places the codebook under measurement projectibility, while Section 7.8 defines taxonomic drift as category-assignment changes 'under a taxonomy or codebook revision'. A codebook revision that redefines a category is thus both a measurement change and a taxonomic change. This is not an empirical limitation of the pilot; it is an internal specification gap in the very typology the framework asks evaluators to adopt. The paper's motivating promise is that the typology tells you which inference license a change retires and what repair is needed; for these changes it does not.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that pass/fail labels in LLM safety evaluation conflate four distinct inference targets, which it calls reference, behavioural, measurement, and taxonomic projectibility. It introduces \"adversarial pragmatics\" as a construct covering instruction conflicts, embedded commands, quotation, scope, deixis, and indirect speech acts, and it releases an 18-item seed benchmark, a schema validator, a 54-row pilot with author labels, and a six-cell LLM-judge validation. The principal claim is methodological: evaluation results should be recorded and interpreted separately for each bearer, with a declared claim register and separate warrants for each projection. The pilot is explicitly framed as a development artifact, not a validated measurement instrument, and the paper is unusually candid about its own limitations, including the absence of independent adjudication, the lack of strict minimal pairs, wide item-clustered intervals, and the circularity of the first judge pass.","tokens_in":34332,"tokens_out":8463,"duration_ms":78985,"significance":"If the framework works, it is a valuable contribution to safety-evaluation methodology: it gives evaluators a vocabulary and an audit trail for separating reference, behaviour, measurement, and taxonomy, and it discourages the improper aggregation of pass/fail labels. The paper is exemplary in its honesty and reproducibility: the repository ships machine-checkable validators, reproducible commands, raw pilot outputs, per-row labels, and a deterministic fake-data calibration pass. The negative LLM-judge results are a useful demonstration that an unvalidated judge can look reasonable in aggregate while missing exactly the minority classes that matter for safety evaluation. The central methodological claim is independent of the pilot's weak empirical power, and the paper does not overclaim what the pilot establishes. The main weakness is a specification gap in the four-bearer typology itself, which needs to be tightened before the framework can deliver its promised unique attribution of inference-license changes.","major_comments":[{"comment":"The four-bearer decomposition is under-specified for the class of revisions it is meant to disambiguate. The reference bearer is defined as the licensed-action relation generated by a \"frozen stipulated-regime template and output-blind adjudication procedure,\" while the measurement bearer is defined to include a \"versioned scorer, codebook, implementation, adjudication procedure, criterion semantics, and enumerated prompt and evidence-access perturbations\" (Section 8). An adjudication-procedure change therefore retires both the reference and the measurement projection, and no stated rule determines whether both must be re-warranted or which record is authoritative. The same overlap occurs for codebooks: Section 1 places the codebook under measurement projectibility, while Section 7.8 defines taxonomic drift as category-assignment changes \"under a taxonomy or codebook revision,\" making a codebook revision simultaneously a measurement and a taxonomic change. Because the central claim is that the typology tells an evaluator \"which bearer a given change moved,\" these overlaps leave the principal inference not uniquely defined. Please add an explicit multi-bearer change rule (e.g., a stated precedence relation, or a requirement that every revision declare and retire all affected bearers and that each retired projection be re-warranted separately).","section":"§1, §7.8, §8"},{"comment":"The pilot's pair-level diagnostics are confounded by design: Table 3 states that every seed pair changes the required response act along with its designated control dimension, so no pair is a strict minimal pair. For example, P001 varies source role but also changes the response act from summarizing to token emission. Consequently, the family-level pass patterns in Figure 4 — \"three pairs passed in all three observed cells\" and the zero-pass pairs — cannot be attributed to the designated control dimension, and the pair-level quantities cannot be read as source sensitivity, scope robustness, or reference stability. The paper is transparent about this, and the statistic is honestly named \"joint pair completion,\" but the discussion in Section 6 and the Figure 4 caption should state explicitly that with the current seed items, pair-level results are descriptive of the specific prompt-and-response configuration only and do not estimate a causal contrast.","section":"§4, §6, Figure 4"}],"minor_comments":[{"comment":"The informal contractions (\"It's\", \"can't\") are out of place in a formal journal style; please rewrite them in full.","section":"Abstract and throughout"},{"comment":"The edge labels (\"fixed output-blind criteria kept separate\", \"route validated per criterion\", etc.) are telegraphic; a fuller caption mapping each edge to the corresponding bearer would make the figure self-contained.","section":"Figure 1"},{"comment":"The \"Status\" column entries \"diagnostic\" and \"referencereview\" should be spelled out or defined in the caption (e.g., \"reference review\" and \"diagnostic contrast\").","section":"Table 3"},{"comment":"The framework relies on two companion manuscripts (Reynolds 2026a, 2026b) for \"typed authorization semantics\" and \"evidentiary assurance\"; since these are not yet available, please add a concise summary of the relevant definitions so the present paper is self-contained on the points where it depends on them.","section":"Section 8"},{"comment":"The entry \"T ensor Trust\" contains a spurious space and should read \"TensorTrust\".","section":"Supplement, Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest, well-scoped, and reproducible; the main risk is the bearer-overlap specification gap, which directly affects the paper's central methodological promise. I would not require new experiments, but I would require the typology to be made precise before acceptance. The heavy citation of the author's own unpublished companion papers is a partial impediment to evaluation; the author should be asked to summarize the needed definitions. The manuscript is otherwise in good shape for a methodological venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper's central move is right: a single pass/fail label in safety evaluation conceals several distinct inference targets, and reporting them separately beats compressing them. The four-bearer decomposition (reference, behavioural, measurement, taxonomic) goes beyond SEP, AgentDojo, IHEval, BIPIA, InjecAgent, and TensorTrust, and the pilot is refreshingly honest about its own circularity.\n\nWhat's actually new: the typed decomposition and the criterion-separated protocol. The pilot data is not the point; the diagnosis is. The judge assessment shows that a rubric-aided judge can look good in aggregate precisely by never emitting the minority class—a useful negative result. The item-clustered intervals and partial pooling are done carefully.\n\nThe main problem is internal. The bearer definitions overlap. An adjudication-procedure change is both a reference-bearer change and a measurement-bearer change (Section 8); a codebook revision is both measurement and taxonomic (Section 1 vs Section 7.8). The paper's motivating promise is that one record can tell you which bearer a change moved, but for this natural class of changes it cannot. That is a specification gap in the typology the framework asks evaluators to adopt, not an empirical limitation of the pilot. It is fixable—add a priority rule or declare which dimension is primary for a given change—but it should be addressed before the framework is used as a discipline.\n\nThe pilot weaknesses are disclosed rather than hidden: author-only labels, self-grading judge, visible answer key, non-minimal pairs. Eighteen clusters is small and the intervals are appropriately wide. Those are not fatal for a methodology paper, but they do mean the empirical contribution is a demonstration of executability, not a validated instrument.\n\nWho gets value: evaluation teams building safety benchmarks and anyone thinking about what a benchmark result licenses. It deserves a serious referee; the likely outcome is revision with a request to tighten the bearer definitions. I would send it out.\n\nBest.","headline":"A genuinely useful methodological proposal for separating what a safety-evaluation label licenses, with an honest but tiny pilot; the main weakness is that the four-bearer typology itself is under-specified for a natural class of changes.","tokens_in":34772,"tokens_out":3025,"would_cite":true,"duration_ms":24988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Safety-evaluation labels should name what they license, not just pass or fail.","keywords":["AI safety evaluation","adversarial pragmatics","prompt injection","instruction hierarchy","projectibility","LLM judges","assessment validity","benchmark diagnosis"],"falsifier":"Run strict minimal-pair versions of the nine contrasts—identical operative string, payload, and expected response act, varying only the declared control dimension—with output-blind independent adjudication, and check whether the pair-level source-sensitivity and judge-failure profiles of the pilot reproduce; if they do not, the per-dimension diagnoses are confounded by the response-act changes.","tokens_in":33720,"feed_emoji":"🛡️","tokens_out":8211,"duration_ms":65673,"temperature":0.7,"pith_summary":"This paper tries to establish that safety-evaluation labels in language-model testing are inference licenses, and that a single pass/fail label collapses four distinct targets a result could speak for: the reference (what a stipulated regime licenses), configured-system behaviour, the interpretation of evaluator outputs, and taxonomic assignment. It introduces a diagnostic framework for adversarial pragmatics—safety-relevant behaviour where instruction status, source authority, quotation, scope, deixis, or indirect speech act must be inferred—and an 18-item seed benchmark built from paired contrasts. A 54-row pilot shows the framework is implementable and exposes methodological limits: a first LLM judge with the expected answer visible still missed safety-relevant minority classes, and item-clustered agreement intervals are too wide to rule out a constant labeller on most families. The paper's intended use is diagnosis of items, rubrics, and evaluator behaviour, not deployment certification, vendor ranking, or a general safety score.","feed_headline":"Pass/fail labels hide four distinct safety-failure modes","feed_subtitle":"A diagnostic framework argues evaluation results should name the inference they license, not just pass or fail.","key_machinery":"The carrying object is the four-way projectibility-bearer decomposition—reference, behavioural, measurement, and taxonomic—applied to every evaluation result, together with an evidence chain that runs from a stipulated regime through licensed reference, model response, criterion-specific judgments, scoring route, interpretation, and projected target. The benchmark operationalizes this through paired contrasts that vary a designated control dimension such as source role, authority, pragmatic status, scope, reference resolution, speech-act force, application surface, or refusal status, using harmless payload tokens. The seed schema records contrast pairs, expected behaviour fixed before outputs are inspected, and separate label fields for task success, policy compliance, safety risk, refusal outcome, failure attribution, and confidence, so that disagreements can be routed to item, protocol, or construct diagnostics rather than averaged away.","core_discovery":"The paper's central claim is that a single evaluation label cannot carry four projectibility bearers at once. Reference projectibility concerns the licensed-action relation fixed by a frozen stipulated-regime template and output-blind adjudication; behavioural projectibility concerns a versioned evaluation-configuration family; measurement projectibility concerns a versioned scorer, codebook, and adjudication procedure; and taxonomic projectibility concerns a versioned taxonomy and assignment procedure. Because these can fail independently and require different repairs, the paper argues that evaluation results should be typed by bearer rather than compressed into one label. It supplies an adversarial-pragmatics taxonomy, an 18-item seed benchmark, an annotation protocol that keeps task success, policy compliance, risk, refusal, attribution, and uncertainty analytically separate, and a pilot demonstrating both implementability and the need for judge validation.","pith_inferences":["If the projectibility-bearer discipline becomes standard, model cards and evaluation reports will need versioned entries not just for the model but for the wrapper family, scorer codebook, adjudication procedure, and taxonomy, so that a result change can be attributed to the component that moved.","The same reasoning transfers to safety-incident review: a transcript should be split into observable failure signature, failure locus, causal attribution, and evidential status before any strong claim about model intent is made; the paper makes this split for its own labels but only sketches it for incidents.","A direct test of the framework is whether reference-concordant response changes in held-out wrappers systematically exceed matched nuisance and placebo changes; that is the paper's own next-stage design, and its success or failure would separate genuine source-sensitivity diagnosis from wrapper-specific artefacts.","A practical extension for judge validation is to require class-specific recall on minority labels, especially middle categories like partial, alongside overall agreement, since the pilot's strongest-looking judge reached its aggregate edge partly by never emitting that category."],"forward_implications":["Evaluation results should be reported and interpreted separately for the reference, behavioural, measurement, and taxonomic bearers, because each can fail while the others hold.","LLM judges must be validated per criterion and per information condition; the pilot shows aggregate agreement can sit at or below a constant-labeller baseline while minority safety-relevant classes such as partial successes and risk-labelled rows are missed entirely.","Pair-level source-sensitivity and robustness metrics require strict minimal pairs with the same expected response act and only the control dimension changed; the current seed pairs are development contrasts whose pair statistics are joint completions, not isolated causal effects.","Any aggregate decision rule—what to collapse, what base rate of unsafe items to assume, and what utility to assign to over-refusal versus under-refusal—is an assessment use requiring its own structural and consequential validation, not a neutral readout.","The benchmark is designed to diagnose items, rubrics, and evaluator disagreement rather than to certify deployments, rank vendors, or produce a general safety score."],"supporting_citations":[{"why":"Contributes the notion of projectibility that the paper's four inference targets are built on.","marker":"Goodman (1955)"},{"why":"Defines the instruction hierarchy and privileged instructions that the authority family of seed items tests.","marker":"Wallace et al. (2024)"},{"why":"Formalizes instruction-data separation, the contrast the embedded-command family operationalizes.","marker":"Zverev et al. (2025)"},{"why":"Establishes indirect prompt injection in LLM-integrated applications, motivating the source-role and embedded-command contrasts.","marker":"Greshake et al. (2023)"},{"why":"Supplies AgentDojo, a dynamic agent-environment benchmark used as a comparison point and as the tool-result surface for the transcript family.","marker":"Debenedetti et al. (2024)"},{"why":"Shows that LLM judges' trained safety priors resist rubric changes, supporting the paper's demand for route-specific judge validation.","marker":"Alloula et al. (2026)"},{"why":"Provides the baseline claim that LLM-as-judge can approximate human preferences in some settings, against which the pilot's judge-failure profile is a caution.","marker":"Zheng et al. (2023)"},{"why":"Supplies the Type-M and Type-S error analysis used to retire the earlier pass/fail gates and to shrink the observed rubric effect.","marker":"Gelman and Carlin (2014)"},{"why":"Frames the validation argument that validity attaches to a proposed interpretation and use of results, which the paper adopts for its four inference targets.","marker":"Messick (1995)"}],"fun_headline_variants":["Four failure modes hide behind a single pass/fail label","Safety labels compress four distinct failure modes into one","Adversarial pragmatics uncovers four failure modes in one label","One pass/fail label hides four separate safety failure modes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's diagnostic power rests on the seed items instantiating the intended pragmatic distinctions closely enough that failures can be attributed to a designated control dimension; the paper itself notes that every seed pair changes the required response act along with the control dimension, so that attribution is not yet isolated.","fun_headline_variants_meta":{"raw":{"variants":["Four failure modes hide behind a single pass/fail label","Safety labels compress four distinct failure modes into one","Adversarial pragmatics uncovers four failure modes in one label","One pass/fail label hides four separate safety failure modes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000533,"raw_usage":{"total_tokens":2605,"prompt_tokens":1025,"completion_tokens":1580,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":1513}},"tokens_in":641,"tokens_out":1580,"duration_ms":8790,"temperature":1.0,"reasoning_tokens":1513,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:35:38.891395+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run strict minimal-pair versions of the nine contrasts—identical operative string, payload, and expected response act, varying only the declared control dimension—with output-blind independent adjudication, and check whether the pair-level source-sensitivity and judge-failure profiles of the pilot reproduce; if they do not, the per-dimension diagnoses are confounded by the response-act changes.","supporting_citations":[{"cited_title":"Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators","cited_arxiv_id":"2606.07874","evidence_quote":"Shows that LLM judges' trained safety priors resist rubric changes, supporting the paper's demand for route-specific judge validation."}],"review_version":3}