{"id":"05d6488e-5cda-471c-b1bb-120bb3446a2e","arxiv_id":"2607.05264","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"On real industrial CCTV, the best of nine VLMs reaches 42.6% action accuracy versus 84.6% human performance, with provenance audits showing up to 17pp inflation from unaudited VLM labels.","lead":"SteelBench is a real steel-plant CCTV benchmark that tests vision-language models on distant workers, dust, steam, glare, and safety rules. Current models top out at 42.6% action accuracy versus 84.6% for humans, and unaudited model-assisted labels can inflate same-family scores by up to 17 points.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-disclosed residual-anchoring concern.","rationale":"The paper's strongest claims are empirical measurements (Table 5, Tables 3–4, CRG definition Eq. 1, DRS checklist) that are directly supported by the released clips, annotations, and inference outputs. The only plausible load-bearing vulnerability is residual expert inheritance of Qwen pre-fill patterns on ambiguous distant scenes; the authors already measure this via the three-level audit, report the 17 pp same-family inflation, and treat full-GT accuracy only as a ranking signal. Because the audit is the contribution rather than an afterthought, and because the reported gaps dwarf the residual contamination rates, the concern does not overturn the claims. The concrete test above is the natural next verification step already enabled by the public blind slice; a null result would leave the ACCEPT verdict intact. I therefore leave the reader's ACCEPT / high-confidence assessment unchanged.","tokens_in":33802,"tokens_out":595,"duration_ms":5915,"concrete_test":"On the released 102-clip blind subset, recompute the 9-model action accuracies and the human reference using only blind Tier-1 labels (no VLM pre-fill) against the same expert reference; if the best-model accuracy rises above ~50% or the human–model gap shrinks below ~25 pp, residual anchoring would be understated and the headline gap would need re-statement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption is correctly identified and is already the paper's own central methodological contribution rather than a hidden soft spot. Level-1 direction analysis (CLR 92.8%, CR 6.9×, harmful anchoring 3.9% on 2,829 field comparisons), Level-2 provenance gradient (blind 37.2% → proper-chain 57.4% → VLM-sourced 77.7% for same-family Qwen), Level-3 human reference (84.6%, κ=0.82 on 370 proper-chain pairs), blind-vs-anchored IAA stability for action/PPE, and exclusion of the rubber-stamping annotator together quantify rather than conceal residual dependence. The 42 pp human–model gap, CRG 0.375–0.582, and ≤2/5 DRS results remain large relative to the measured 3.9% residual and the ~6.6 pp universal anchored-GT inflation. Single-facility scope is disclosed and partially mitigated by 6-domain internal diversity (8–21 pp spreads). No additional load-bearing inconsistency is required for the central claims to hold within the stated diagnostic scope.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"SteelBench is a diagnostic benchmark for vision-language models on real industrial CCTV from an operational integrated steel plant. It provides 1,345 densely annotated clips (from 149 hours / 10,024 candidates) with per-worker actions (25 classes), PPE, spatial context, visibility, and safety-rule labels under a 2-layer schema. The central methodological contribution is a three-level provenance-aware audit of model-assisted annotation (label influence via CLR/CR and anchoring bias; GT provenance sensitivity across blind / proper-chain / VLM-sourced slices; human reference). Evaluating nine VLMs, the paper reports a large human–model gap (best 42.6% action accuracy vs 84.6% human reference), up to ~17–40pp inflation under VLM-sourced GT for same-family evaluation, high compositional reasoning gaps (CRG 0.375–0.582 even on action-correct instances), fragmented robustness/calibration, and at most 2/5 DRS diagnostic checks passed. Ablations cover prompt variants, frame density, threshold sensitivity, and domain/site/condition breakdowns; data and code are released.","tokens_in":34214,"tokens_out":1291,"duration_ms":18799,"significance":"If the results hold, the paper is a substantial contribution on two fronts: (i) a realistic industrial surveillance evaluation surface that existing consumer, egocentric, and simulated industrial benchmarks do not provide, and (ii) an explicit, measurable protocol for auditing circularity in model-assisted benchmark construction—an issue widely acknowledged but rarely quantified. Strengths include multi-slice provenance experiments, exclusion of a rubber-stamping annotator via audit signals, human reference on 370 person-level pairs with κ, bootstrap CIs, prompt/frame ablations, DRS threshold sensitivity, and public release under CC-BY-NC-4.0 with code. The diagnostic framing (recognition vs robustness vs calibration vs safety reasoning) is more useful for deployment decisions than a single leaderboard score. Single-facility scope and residual anchoring are real limits but are disclosed and partially mitigated by internal domain diversity and quantified residual rates.","major_comments":[{"comment":"§5.1 Level 3 and Appendix D.7: The headline 84.6% human reference (and thus the 42pp gap) is measured on proper-chain pairs where Tier-1 annotators saw Qwen3-VL-235B pre-fills before expert verification. Blind-condition human accuracy against the expert is described as lower but is not reported as a primary number alongside 84.6%. Because residual harmful anchoring is 3.9% and the paper’s own Level-2 gradient shows large provenance effects, the main human–model gap claim would be stronger if blind human accuracy (and its n) were stated in §5.1 / Table 4 as a co-primary reference, with proper-chain retained as the operational annotation-workflow baseline.","section":null},{"comment":"§4 and Appendix C.2.6–C.2.8 (DRS / DWA): DWA’s taxonomic distances (0 / 0.33 / 0.60 / 0.70–1.00) and the five DRS thresholds are load-bearing for the claim that no model is deployment-ready (≤2/5 checks). Threshold sensitivity (±10%) is analyzed and margins for CRG and DWA are large, which is good; however, the distance weights themselves are not justified against plant incident severity or officer ranking beyond narrative rationale. A short sensitivity table over alternative distance schemes (or collapsing to group-level accuracy) would show whether the DWA fail is robust or scheme-dependent, analogous to the existing threshold ablation.","section":null}],"minor_comments":[{"comment":"Abstract vs §5.1: Abstract says unaudited VLM-sourced GT can inflate same-family accuracy by “up to 17 percentage points,” while Table 4 / text report 37.2% → 57.4% → 77.7% (larger spans). Align the abstract figure with the exact contrast intended (e.g., proper-chain vs VLM-sourced ≈20pp, or blind vs VLM-sourced).","section":null},{"comment":"Figure 4 / Table 6: Classes C4 and D3 have n<15 and are correctly excluded from per-class claims, but the heatmap still displays them without a clear visual marker; add a hatch or footnote so readers do not over-read those cells.","section":null},{"comment":"Table 1 vs Appendix A.3: Layer-2 / Layer-1 clip counts differ slightly across places (805/540 vs 807/538). Harmonize final counts after annotator_10 exclusion.","section":null},{"comment":"§2 and Figure 2: “Class balancing” in the abstract/curation text sits awkwardly next to the retained natural long-tail distribution; clarify that stratified sampling enforces minimum support where available rather than uniform class balance.","section":null},{"comment":"Appendix E.2 V2: Free-form descriptions are mapped by GPT-4o-mini; state mapper agreement or a small human audit so V2 accuracy is not over-interpreted relative to V1/V3.","section":null},{"comment":"Minor polish: arXiv id / preprint date consistency; expand first use of MAI/MAC/SA in the main text before Figure 1; fix occasional spacing (e.g., “CCTV ,”, “A V A”).","section":null}],"recommendation":"minor_revision","confidential_remarks":"Strong fit for a datasets-and-benchmarks or applied CV venue. The provenance audit is the distinctive methodological contribution and is executed carefully enough that residual circularity is a quantified limitation rather than a hidden flaw. I would not block on single-facility scope given the disclosed 6-domain internal spreads and planned expansion. No concerns about ethics disclosure or data release beyond the usual NC license tradeoff for industrial footage."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a solid datasets-and-benchmarks paper. The new pieces are real operational CCTV (not sim), dense per-worker labels (action + PPE + spatial + safety), and an explicit three-level provenance audit that measures how VLM pre-fills move labels and scores. That combination is not in Kinetics-style sets, IndustryEQA, or PPE-only safety datasets.\n\nWhat they do well is the audit itself. Blind / proper-chain / VLM-sourced slices, productive vs harmful anchoring (CLR 92.8%, CR 6.9×, harmful 3.9%), exclusion of a rubber-stamping annotator, human reference 84.6% (κ=0.82 on 370 pairs), bootstrap CIs, prompt and frame ablations, and public HF + code. Same-family Qwen inflation up to ~17pp (and ~6.6pp universal anchored-GT lift) is the result people will cite. The capability fragmentation is also clean: best action accuracy 42.6%, CRG 0.375–0.582 even on correct actions, no model >2/5 DRS checks, false-alarm vs false-safe split. Math and metrics are straightforward; citations cover the right prior work without padding.\n\nSoft spots are the ones they already flag. Residual 3.9% harmful anchoring means the expert reference is not fully independent of Qwen pre-fills on ambiguous distant scenes—so the human–model gap and inflation numbers could be slightly biased. Single facility is real, though 6 internal domains with 8–21pp spreads help. DRS thresholds are plant-officer-derived and free parameters; they show sensitivity. Groups C/D are thin at class level. None of that overturns the central claims: the gaps stay large relative to the measured residual.\n\nThis is for people building or evaluating industrial VLMs, safety monitoring, or model-assisted annotation pipelines. Not a theory paper. I would bring it to reading group, cite the provenance numbers and the real-footage gap, and send it to peer review. Accept with the usual requests for clearer residual-bias bounds and multi-site plans.","headline":"Real-plant industrial VLM benchmark with a quantified provenance audit; the 42pp human gap and 17pp same-family inflation are the numbers that matter.","tokens_in":34883,"tokens_out":544,"would_cite":true,"duration_ms":6011,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"On real steel-plant CCTV, the best vision-language model hits only 42.6% action accuracy against an 84.6% human reference, and correct actions still produce wrong safety judgments 37–58% of the time.","keywords":["vision-language models","industrial surveillance","action recognition","safety reasoning","annotation provenance","benchmark","PPE compliance","CCTV"],"falsifier":"Re-label a blind, model-free subset of the same clips with independent experts and re-run the nine models: if the human–model gap collapses or the provenance inflation shrinks dramatically, the central performance claim does not hold.","tokens_in":34696,"feed_emoji":"🏭","tokens_out":661,"duration_ms":5549,"temperature":0.7,"pith_summary":"SteelBench argues that existing video benchmarks do not test vision-language models under the conditions of real industrial surveillance: distant workers, dust, steam, glare, occlusion, and overlapping activities. The authors release 1,345 densely annotated clips from an operating steel plant, each carrying per-worker actions, PPE status, spatial context, and safety-rule labels. Because the labels themselves were built with model help, they also introduce a provenance-aware audit that measures how much VLM pre-fills shape final ground truth. The audit shows that unaudited same-family labels can inflate measured accuracy by up to 17 percentage points. Across nine models, recognition, robustness, calibration, and safety reasoning fail independently, and no model passes more than two of five deployment-readiness checks. The paper’s claim is that reliable industrial activity understanding requires provenance-aware, failure-mode-specific evaluation rather than a single accuracy leaderboard.","feed_headline":"Best VLM scores 42.6% on real plant CCTV vs 84.6% human","feed_subtitle":"Even correct actions yield wrong safety calls 37–58% of the time; provenance can inflate scores 17 points","key_machinery":"The provenance-aware audit protocol: a three-level procedure that measures label influence from VLM pre-fills (productive vs harmful anchors and overrides), re-scores the same model predictions against blind, proper-chain, and VLM-sourced ground truth, and reports a human reference from expert-reviewed labels.","core_discovery":"On real operational CCTV, the strongest of nine vision-language models reaches only 42.6% action accuracy versus an 84.6% human reference; even when the action is correct, 37–58% of safety judgments remain wrong; and unaudited VLM-sourced ground truth can inflate same-family accuracy by up to 17 percentage points. No model passes more than two of five diagnostic deployment checks.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Best of nine VLMs reaches 42.6% action accuracy on real plant CCTV vs 84.6% human","Even correct actions yield wrong safety calls 37–58% of the time on SteelBench","Unaudited VLM labels can inflate same-family accuracy by up to 17 points","No model passes more than two of five industrial CCTV diagnostic checks","SteelBench: VLMs lag humans sharply on distant, occluded plant surveillance"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The expert-verified labels used as the gold standard remain independent enough of the original model pre-fills that the large human–model gap and the reported accuracy inflation are trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Best of nine VLMs reaches 42.6% action accuracy on real plant CCTV vs 84.6% human","Even correct actions yield wrong safety calls 37–58% of the time on SteelBench","Unaudited VLM labels can inflate same-family accuracy by up to 17 points","No model passes more than two of five industrial CCTV diagnostic checks","SteelBench: VLMs lag humans sharply on distant, occluded plant surveillance"]},"model":"grok-4.5","effort":"low","cost_usd":0.005856,"raw_usage":{"total_tokens":1592,"prompt_tokens":876,"num_sources_used":0,"completion_tokens":117,"cost_in_usd_ticks":58560000,"prompt_tokens_details":{"text_tokens":876,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":599,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":876,"tokens_out":117,"duration_ms":4746,"temperature":1.0,"reasoning_tokens":599,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T07:27:00.241756+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-label a blind, model-free subset of the same clips with independent experts and re-run the nine models: if the human–model gap collapses or the provenance inflation shrinks dramatically, the central performance claim does not hold.","supporting_citations":[],"review_version":2}