{"id":"99f23e91-665e-47f1-b878-c6984e9ce247","arxiv_id":"2608.10673","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A marker-based review of 83 Compton imaging papers shows the field reports single operating points and omits latency, cost and source-complexity terms, and that deployment commitment tracks gap-filling and handover count rather than publication volume.","lead":"This paper reviews 83 Compton imaging papers and finds that the field reports performance at a single operating point, while transfer between applications requires gradients like count-rate sweeps. It also finds that deployment commitment tracks whether the incumbent can do the task and how many handovers the output must cross, rather than the size of the literature or stated need.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Table 5 'Committed' measure is defined per domain (flight selection, clinical endpoint, field measurement), so the claimed inverse-of-publication ordering may reflect unequal milestone strictness rather than deployment commitment; a uniform recoding is needed.","rationale":"The paper's central descriptive counts are carefully documented: denominators are stated, corrections are conservative, and the authors score their own work alongside the rest of the corpus. The 65/83 and zero-sweep findings, if the coding is reliable, are a real and important observation about the field's reporting conventions. The weakest point in the central claim is the cross-domain comparison, exactly where the reader placed it. The per-domain definitions of Committed are not interchangeable: 'clinical endpoint' is a later and rarer milestone than 'field measurement' with a prototype, while 'selected mission' is programmatic and may never fly. Because the inverse-of-publication sentence is one of the paper's headline findings and is meant to be explained by incumbent gap and boundary count, a spurious ordering from threshold differences would weaken the novelty claim. The proposed check settles this by holding the definition fixed across domains. I do not move the verdict to CONDITIONAL because the paper already labels the domain finding an observed regularity rather than a law, the main metrological argument is independent of it, and the limitations are reported visibly. The concern is a calibration issue in a secondary claim, not a soundness failure of the review's primary counts.","tokens_in":26781,"tokens_out":11245,"duration_ms":115875,"concrete_test":"Re-code Table 5 with one uniform commitment definition applied to all five domains, counting for each domain works that show (a) measurements in the operational environment, (b) a formal programmatic selection or dedicated funding line, or (c) a regulatory or clinical-study endpoint. Recompute the ordering of domains and the inverse-of-publication relation under each of these three uniform definitions. If the ranking by incumbent gap and boundary count, and the inverse relation, survive under a definition that does not embed different per-domain thresholds, the concern is refuted; if the ranking changes, the domain-ordering claim should be downgraded to hypothesis.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The domain-ordering claim (Section 5.1 and Conclusions) rests on Table 5's Committed column, whose definition changes per row: astrophysics counts flight hardware, a balloon campaign or a selected mission; particle therapy counts a clinical endpoint; decommissioning and survey count field, on-site or vehicle-borne measurement. These milestones sit at very different points of a deployment pipeline, so the column is not a single quantity. A clinical endpoint is a far stricter and later milestone than a prototype used on site, and a selected mission has not flown. The ordering (decommissioning 72%, astrophysics 59%, particle therapy 6%) and the accompanying statement that 'the distribution of publication is close to the inverse of the distribution of deployment commitment' could therefore be produced by the thresholds chosen rather than by an underlying property of the domains. The paper discloses the definitions and calls the result a regularity, but disclosure does not establish commensurability. The boundary-count ordering is also not strictly monotone in the commitment column (astrophysics, 0 boundaries, 59%; decommissioning, 1 boundary, 72%), so 'orders the last column' overstates the support. The reporting-metrology counts (65/83 at one operating point; count rate swept in 0) are independent of this and remain the paper's stronger evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a meta-review of the Compton imaging literature. It first develops a framework for evaluating imaging chains (four founding questions Q1–Q4, bounded-latency composition, a viability bracket Qmax ≥ Qmin inside Twindow), then reports quantitative counts over 83 full texts: performance reported at one operating point in 65, count rate swept in none, no reported cost of precomputation or model inference time, and sparse reporting of memory, separability, and calibration. It then compares five application domains on a common template (Table 5) and claims that neither the size of the receiving literature, nor the incumbent, nor the stated need orders domains by deployment commitment, whereas two problem-side properties (whether the incumbent can serve the task, and the number of domain boundaries crossed) do so, and that the distribution of publication is close to the inverse of the distribution of deployment commitment. The paper closes with a metrology agenda and a list of quantities that future reports should contain.","tokens_in":26979,"tokens_out":7710,"duration_ms":67813,"significance":"If the reported counts are reliable, the paper documents a substantial and actionable reporting gap: the field publishes stage-level improvements whose composition into an application cannot be evaluated from the text. Its strengths are real: the corpus is enumerated, denominators are stated for each count, markers are defined and applied to full texts in context, the direction of coding corrections is uniform, the authors include their own companion work in the corpus, and the supplementary material allows recomputation. The central descriptive claims about operating-point reporting and missing cost quantities are supported by this transparent accounting, and the proposed metrology remedy is constructive. The domain-ordering claim is more fragile; it rests on a deployment-commitment measure defined differently in each domain and on a small number of data points, and it needs strengthening or qualification before the conclusions can be accepted at their current strength.","major_comments":[{"comment":"The headline statistic is 'Performance is reported at one operating point in 65 of 83 full texts' (Abstract and §7), but §5.3 states that '70 yielded text clean enough for the operating-point and reporting markers.' If the operating-point markers were scored only on those 70 texts, the denominator should be 70, not 83, or the paper must explain how the remaining 13 texts were scored. Since this statistic is the paper's most prominent evidence of the reporting gap, please state the exact denominator for every reported count and reconcile the abstract, §5.3, and the conclusions.","section":"Abstract; §5.3; §7 first bullet"},{"comment":"The domain-ordering claim rests on the 'Committed' column, whose operational definition changes per row: for astrophysics it counts flight hardware, a balloon campaign, or a selected mission; for particle therapy it counts a clinical endpoint; for decommissioning and survey it counts field, on-site, or vehicle-borne measurement. A clinical endpoint is substantially later and stricter than a prototype used on site, and a selected mission has not flown, so the column does not measure a single quantity. The paper discloses these differences and calls the result a regularity, but disclosure does not establish commensurability. Please provide a sensitivity analysis under a common milestone definition (for example, any documented use in an operational setting with a decision taken on the result) or explicitly downgrade the ordering claim to a coding-dependent observation. Relatedly, the boundary-count statement 'That count orders the last column' is not strictly supported: Table 5 gives boundary counts 1 (decommissioning, 72%), 0 (astrophysics, 59%), and 3 (particle therapy, 6%), which is non-monotone; with only three committed domains, 'consistently' overstates the evidence.","section":"Table 5; §5.1; §7 third bullet"},{"comment":"The claim that 'the distribution of publication is close to the inverse of the distribution of deployment commitment' is presented without a quantitative measure of the closeness or a statement of which publication distribution is being used. The 'Ours' column in Table 5 (99, 219, 128, 111, 128) is not inversely ordered across the three domains with commitment values, whereas the declared-application counts in §5.1 (21, 20, 3, 2, 2) are more consistent with the claim. Please specify the exact comparison being made and, if possible, quantify it; as written, the conclusion is too strong relative to the evidence shown.","section":"§5.1; Table 5; §7 fourth bullet"}],"minor_comments":[{"comment":"The text says extended or distributed activity appears in 8 works of 83, while Fig. 3 reports 'extended source 10' on the same denominator; please reconcile these numbers.","section":"§4.1; Fig. 3"},{"comment":"The text says learned models appear in 27 works for image formation, while Fig. 4 reports 'learned model used 31' for the corpus; please clarify whether these are different markers and, if so, state both definitions.","section":"§3.2.3; Fig. 4"},{"comment":"The caption describes Qmax for low- and high-entropy sources and the viable/not viable regions, but the figure itself would benefit from explicit axis labels and a legend identifying the curves; currently the reader must infer which curve is which.","section":"Fig. 2"},{"comment":"Several non-ASCII names are corrupted in the bibliography, for example 'Jelnek' [7], 'Koodziej' [62], and 'Mller' [74]; these should be corrected to Jelínek, Kołodziej, and Müller.","section":"References"},{"comment":"Fig. 1 is reproduced from the authors' companion work [3], which is listed as a 2026 preprint and may not be publicly available; please confirm that permission or a permanent identifier is provided.","section":"Fig. 1; References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful and largely transparent meta-review, and the core reporting-gap finding is worth publishing. The main risk is the domain-ordering section: the deployment-commitment measure is defined differently per domain, the boundary-count ordering is non-monotone, and the 'inverse of publication' claim is not quantified. These issues are fixable within the manuscript's scope by recoding or weakening the claims. The denominator inconsistency in the headline '65 of 83' statistic must also be resolved before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Albiol et al. have written a genuinely useful meta-review: it counts what the Compton imaging literature actually reports, in context, with denominators stated and the authors' own work included in the corpus. The headline numbers are new and robust: 65 of 83 full texts report performance at a single operating point, count rate is swept in none, and no work states precomputation or inference cost. The value-versus-gradient framing is a real contribution, and the stage-by-stage walk of the chain is careful and fair. The bibliometric counts are properly flagged as weaker.\n\nThe soft spot is exactly where the stress-test note puts it. Table 5's 'Committed' column is defined differently per domain—flight hardware or mission selection for astrophysics, a clinical endpoint for particle therapy, field or vehicle-borne measurement for decommissioning. Those milestones sit at very different points of a deployment pipeline, so the column is not one quantity. The claimed ordering, and the inverse-publication statement, may be an artifact of threshold choice. The paper discloses the definitions and calls the result a regularity, but disclosure does not establish commensurability. The boundary-count ordering is also not strictly monotone, as the note observes. This weakens the domain-ordering section but does not touch the reporting-metrology core.\n\nThe manual coding without inter-rater reliability is a limitation, but the authors go further than most: marker definitions and a per-work screening matrix are promised as supplementary data, and the uniform direction of corrections (every one reduced a count) argues for conservatism. Including their own companion work in the corpus is a point in their favor.\n\nThis paper deserves a serious referee. I'd send it to peer review; a good referee should ask for a clearer separation between the robust reporting counts and the softer domain-ordering claim, and ideally a recoding of 'Committed' under a uniform definition. I'd cite it if I worked in this area.","headline":"A transparent and useful meta-review whose reporting-metrology counts are strong; the domain-ordering claim is real but rests on an incommensurable 'Committed' measure and should be softened.","tokens_in":27553,"tokens_out":3325,"would_cite":true,"duration_ms":29336,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 83 Compton-imaging full texts finds performance reported at a single operating point in 65 of them, with count rate swept in none — the quantities transfer requires are the ones the literature omits.","keywords":["Compton camera","gamma-ray imaging","bounded latency","real-time systems","image reconstruction","technology transfer","deployment commitment","operating point"],"falsifier":"Re-code the five application domains with a single commitment metric defined identically across all of them — for instance, a budgeted deployment programme with a committed operator and a stated acceptance test — and re-test whether the ordering by incumbent gap and domain-boundary count survives; the paper itself notes the commitment definitions differ by domain, so this is the direct check. A second, separable test: take any Compton camera, sweep count rate while holding the reconstruction fixed, and publish the resulting latency–quality curve; the paper's finding of zero such sweeps among 83 full texts would be falsified by a single counterexample, and the curve itself is the gradient the paper argues the field is missing.","tokens_in":26561,"feed_emoji":"📷","tokens_out":16362,"duration_ms":131484,"temperature":0.7,"pith_summary":"The paper sets out to determine what evidence a report must contain for a Compton-imaging result to be carried from the stage that produced it to an application that did not build it. Walking the whole chain — detection, digitisation, calibration, event building, reconstruction, and delivery of a result to whoever acts on it — the review scores 83 full texts in context and finds performance reported at a single operating point in 65 of them, with count rate swept in none and no work stating a memory footprint, a precomputation cost, or a learned model's inference time. Measured on one template across five application domains, deployment commitment is ordered not by the size of the receiving literature, the incumbent, or the stated need, but by two properties of the problem: whether the incumbent can serve the task at all, and how many domain boundaries the output must cross before anyone acts — with publication distributed close to the inverse of commitment. The constructive claim is that a defined reporting vocabulary — an operating point, a declared cost, a latency bound per state, a named consumer — turns the viability test $Q_{\\max}(T_{\\mathrm{window}}, H_{\\mathrm{source}}, R_{\\mathrm{acc}}, L_{\\mathrm{total}}) \\ge Q_{\\min}$ into a computation any third party can perform, and that a shared reference object is the one move that requires nobody's permission.","feed_headline":"65 of 83 Compton-imaging papers report one operating point","feed_subtitle":"The field reports values where transfer requires gradients; deployment lands where incumbents can't reach the task.","key_machinery":"The argument is carried by a compositional model of the imaging chain, organised by the four founding questions — what decision the image supports, the window in which it must be available, the quality below which the decision fails, and the extent and complexity of the source — and by the viability bracket that combines them: an application is viable where the maximum attainable quality inside the available window, $Q_{\\max}(T_{\\mathrm{window}}, H_{\\mathrm{source}}, R_{\\mathrm{acc}}, L_{\\mathrm{total}})$, clears the minimum required $Q_{\\min}$. Every stage-level quantity is placed inside that bracket so it can be declared in advance and verified afterwards: bounded latency as a sum of stage bounds $L_{\\mathrm{total}} \\le L_{\\mathrm{transport}} + L_{\\mathrm{order}} + L_{\\mathrm{group}} + L_{\\mathrm{pair}} + L_{\\mathrm{queue}} + L_{\\mathrm{encode}} + L_{\\mathrm{recon}}$; the statistical wait $N/R_{\\mathrm{acc}}$ kept apart from the computational wait, with streaming's gain being $C(N) - L_{\\mathrm{total}}$; and source complexity normalised by $M_{\\mathrm{source}} \\approx (D/\\delta)^d$, so that quoted event counts are comparable only after division by $M_{\\mathrm{source}}$. The review uses this machinery as a scoring scheme — for each stage, what binds first and whether the published text contains the quantity that would compose — which is what produces the counts, and it drives the domain comparison through a second instrument, the one-template measurement of five application domains in Table 5.","core_discovery":"The central finding is a measurement of a reporting convention: performance is reported at one operating point in 65 of the 83 full texts, count rate is swept in none, source complexity is never swept, and the cost of a precomputation or the inference time of a learned model is reported in none. The consequence the paper draws is that superiority over a predecessor and sufficiency for an application are independent statements that accumulate separately: a programme can advance genuinely in the first frame while the second does not move, without any want of rigour. Across five application domains measured on one template, neither the size of the receiving literature, nor of the incumbent, nor of the stated need orders the domains as deployment commitment does, while whether the incumbent can serve the task at all, and how many domain boundaries the output must cross, do so consistently; the distribution of publication is close to the inverse of the distribution of deployment commitment. The positive thesis is that the gap is closable by reporting rather than by more physics: a reference object, a stated operating point and a declared cost make results commensurable, because the quantities the review assembles — bounded latency as a sum of stage bounds, statistical wait $N/R_{\\mathrm{acc}}$ separated from computational wait, event counts normalised by source complexity $M_{\\mathrm{source}} \\approx (D/\\delta)^d$ — are all declarable on paper before a prototype exists.","pith_inferences":["An editor's extension: the same context-scored marker scheme could be applied to neighbouring single-photon imaging literatures, such as coded-mask and collimated gamma cameras; if the single-operating-point convention recurs there, it is a property of stage-organised reporting in imaging generally rather than of Compton cameras specifically.","An editor's extension: the two problem properties that ordered the five domains could serve as a cheap screening test for any proposed new application — count the domain boundaries the output must cross and check whether the incumbent can reach the task — before any detector is committed; the paper identifies the regularity but does not propose it as a predictive tool.","An editor's extension: the paper's reporting vocabulary is directly testable as an intercomparison protocol, a round-robin in which each group images the same reference object and reports a stated operating point, a declared cost and a latency bound per state; the paper names the metrology gap but does not design the exercise.","An editor's extension: the batch-versus-streaming tradeoff implies streaming's benefit concentrates where the statistical wait is longest, which a group that owns a camera could verify by measuring completion time at two event rates — a test the paper's own corpus shows has never been run."],"forward_implications":["Stage-reported improvements, even when genuine, do not compose: a speed-up ratio against a stage baseline cannot be added to a budget, compared with a task window, or used to judge whether an instrument serves an application, so superiority and sufficiency must be tracked separately.","Event counts from different works are comparable only after normalisation by source complexity $M_{\\mathrm{source}}$; a result demonstrated on a point source does not extend to a distributed source by adding events or machines, because conditioning degrades as well as cost.","Adoption is readable in advance from problem properties rather than from market size: whether the incumbent can serve the task at all, and how many domain boundaries the output must cross, ordered the five measured domains exactly as deployment commitment did.","The quantities a report must contain for transfer — an operating point, a cost per resource, a latency bound per state, a named consumer — are all declarable before a prototype exists, and defining a shared reference object is the one step that carries no regulatory or clinical risk.","The statistical wait, not the algorithm, is what binds: since $T_{\\mathrm{batch}} \\approx N/R_{\\mathrm{acc}} + C(N)$ and $T_{\\mathrm{stream}} \\approx N/R_{\\mathrm{acc}} + L_{\\mathrm{total}}$, the algorithmic cost the field publishes on is a diminishing part of the problem, while the physical term it measures least decides whether these instruments reach the applications they invoke."],"supporting_citations":[{"why":"The stage-internal review of Compton image reconstruction, re-read in the paper to document that speed is reported as a relative speed-up with no absolute reconstruction time, so a ratio composes with nothing.","marker":"[1]"},{"why":"The authoritative in-vivo range-verification review whose conclusions — no routine solution exists, no single technique suffices — the corpus cites without engaging, anchoring the claim that the receiving discipline's own statements are available and unused.","marker":"[2]"},{"why":"The 4D prompt-gamma verification study whose hours-long Monte Carlo reference against a five-minute target shows the instrument delivers a comparison carrying a second model, not a quantity.","marker":"[25]"},{"why":"The range-verification feasibility statement that converts an aspiration into a factor (15–40 times required setup-efficiency increase), the only corpus statement answering the source-scaling question in the required form.","marker":"[27]"},{"why":"The clinical study stating the decision ('relevant or non-relevant treatment deviation per field') with a false-positive bound, the corpus's most complete statement of what the image is for and when it must arrive.","marker":"[26]"},{"why":"The essay on essential versus accidental complexity used to argue that a learned model cannot make a cone stop being a surface, and that an inference time never stated exchanges a characterised difficulty for an uncharacterised one.","marker":"[15]"},{"why":"Algebraic spatial sampling as the representation that pays the per-event conical surface once, evidence that a different lower bound comes from a change of representation rather than a faster implementation.","marker":"[16]"},{"why":"The 75 frames-per-second production rate reported with a delivery-accuracy criterion, the corpus's clearest case of a number attached to the consumer rather than to the computation.","marker":"[17]"},{"why":"The authors' own bounded-latency spherical-histogram reconstruction, included in the corpus and scored by the same markers, providing the representation-change example and the handover-magnitude terms.","marker":"[3]"},{"why":"The argument that a new party's cost to an existing workflow is coordination, not work, used to measure adoption by the steps a device adds to a schedule.","marker":"[20]"}],"fun_headline_variants":["65 of 83 Compton papers: one operating point, zero sweeps","Compton imaging reports one dot, not a gradient","Publication order is inverse to deployment commitment","65/83 papers: no sweep, no cost, no transferable result","One point per paper: why Compton results don't transfer"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The domain-ordering claim assumes the five 'Committed' measurements in Table 5 are commensurable even though each is defined differently per domain — flight hardware or a selected mission for astrophysics, a clinical endpoint for particle therapy, field or vehicle-borne measurement for decommissioning — so if those definitions do not measure the same underlying quantity, the consistent ordering by incumbent gap and boundary count could be an artifact of the coding rather than a property of the application domains.","fun_headline_variants_meta":{"raw":{"variants":["65 of 83 Compton papers: one operating point, zero sweeps","Compton imaging reports one dot, not a gradient","Publication order is inverse to deployment commitment","65/83 papers: no sweep, no cost, no transferable result","One point per paper: why Compton results don't transfer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00033,"raw_usage":{"total_tokens":1941,"prompt_tokens":1151,"completion_tokens":790,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":767,"completion_tokens_details":{"reasoning_tokens":707}},"tokens_in":767,"tokens_out":790,"duration_ms":7654,"temperature":1.0,"reasoning_tokens":707,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:39:43.094650+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code the five application domains with a single commitment metric defined identically across all of them — for instance, a budgeted deployment programme with a committed operator and a stated acceptance test — and re-test whether the ordering by incumbent gap and domain-boundary count survives; the paper itself notes the commitment definitions differ by domain, so this is the direct check. A second, separable test: take any Compton camera, sweep count rate while holding the reconstruction fixed, and publish the resulting latency–quality curve; the paper's finding of zero such sweeps among 83 full texts would be falsified by a single counterexample, and the curve itself is the gradient the paper argues the field is missing.","supporting_citations":[{"cited_title":"Feasibility of in-vivo 4D prompt-gamma treatment verification in proton therapy for pancreatic cancer.Physics and Imaging in Radiation Oncology, 40:101027, 2026","cited_arxiv_id":null,"evidence_quote":"The 4D prompt-gamma verification study whose hours-long Monte Carlo reference against a five-minute target shows the instrument delivers a comparison carrying a second model, not a quantity."},{"cited_title":"Range verification by means of prompt-gamma detection in particle therapy.Springer, 2020","cited_arxiv_id":null,"evidence_quote":"The range-verification feasibility statement that converts an aspiration into a factor (15–40 times required setup-efficiency increase), the only corpus statement answering the source-scaling question in the required form."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The clinical study stating the decision ('relevant or non-relevant treatment deviation per field') with a false-positive bound, the corpus's most complete statement of what the image is for and when it must arrive."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The essay on essential versus accidental complexity used to argue that a learned model cannot make a cone stop being a surface, and that an inference time never stated exchanges a characterised difficulty for an uncharacterised one."},{"cited_title":"Study of 3D fast Compton camera image reconstruction method by algebraic spatial sampling.Nuclear Instruments and Methods in Physics Research A, 954:161345, 2020","cited_arxiv_id":null,"evidence_quote":"Algebraic spatial sampling as the representation that pays the per-event conical surface once, evidence that a different lower bound comes from a change of representation rather than a faster implementation."},{"cited_title":"Real-time tracking of the bragg peak during proton therapy via 3d protoacoustic imaging in a clinical scenario.npj Imaging, 2(1), 2024","cited_arxiv_id":null,"evidence_quote":"The 75 frames-per-second production rate reported with a delivery-accuracy criterion, the corpus's clearest case of a number attached to the consumer rather than to the computation."},{"cited_title":"Brooks.The Mythical Man-Month: Essays on Software Engineering","cited_arxiv_id":null,"evidence_quote":"The argument that a new party's cost to an existing workflow is coordination, not work, used to measure adoption by the steps a device adds to a schedule."}],"review_version":1}