{"id":"967afd9c-1243-41f6-aabb-cb4f1c319d43","arxiv_id":"2606.19390","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Presents a framework for generating cryptographically signed CSAF VEX advisories for agentic AI by binding SBOM/AIBOM artifacts to deterministic environment capture and runtime telemetry.","lead":"The paper describes a framework linking SBOM and AIBOM artifacts with runtime telemetry to automatically generate signed CSAF VEX security advisories for agentic AI systems, validated via deterministic replay. A smart generalist might read it to understand emerging methods for automating security in autonomous AI deployments.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Synthetic workloads (50-5000 components) plus public vuln datasets may not capture real activation conditions or yield reliable exploitability signals.","rationale":"The reader's weakest_assumption matches the load-bearing point exactly. Because the evaluation description is limited to synthetic data and the full text was unavailable to the reader, the concern remains unaddressed; no adjustment to UNVERDICTED is warranted.","tokens_in":1624,"tokens_out":300,"duration_ms":16673,"concrete_test":"Take one of the paper's 50-5000 component synthetic workloads, inject 5-10 known exploitable conditions with ground-truth labels, run the full pipeline, and measure precision/recall of the generated VEX statements; repeat on an equivalent real deployed agentic system if available.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The framework's core claim is that exploitability can be computed from declared artefacts, observed activation conditions, and enforced policies, then turned into signed CSAF-VEX advisories validated by deterministic replay. This computation is evaluated only on ~10000 synthetic component entries drawn from 50-5000-component Agentic AI workloads plus OSV/GitHub/KEV/EPSS data. The mapping from synthetic telemetry to real-world false-positive/negative rates for exploitability is therefore the least-secured link; nothing in the described evaluation demonstrates that the static+runtime evidence combination generalizes beyond the synthetic regime.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a protocol-driven framework that binds SBOM and AIBOM artefacts to deterministic environment capture and structured runtime telemetry. Exploitability is computed from declared artefacts, observed activation conditions, and enforced execution policies. CSAF-VEX advisories are generated from the combined static and runtime evidence, cryptographically signed, and validated through deterministic replay. The evaluation uses approximately 10000 component entries drawn from synthetic Agentic AI workloads of 50 to 5000 components, incorporating OSV, GitHub Advisory, KEV, and EPSS datasets.","tokens_in":1741,"tokens_out":426,"duration_ms":24053,"significance":"If the framework's claims hold, it would provide a reproducible, evidence-based method for generating execution-bound security advisories tailored to agentic AI systems, potentially improving vulnerability management by linking static declarations with runtime observations and adding cryptographic verifiability. The deterministic replay validation is a notable strength for reproducibility. However, the significance is constrained by the exclusive use of synthetic workloads, which leaves open whether the exploitability signals generalize to real-world conditions.","major_comments":[{"comment":"Evaluation description (abstract and evaluation section): The manuscript states that the evaluation uses approximately 10000 component entries but provides no details on computation methods for exploitability, specific results, error handling, or validation outcomes. This absence makes it impossible to assess whether the static-plus-runtime evidence combination supports the central claim of reliable exploitability computation from artefacts, activation conditions, and policies.","section":"Evaluation description"},{"comment":"Evaluation description (abstract and evaluation section): The framework is tested exclusively on synthetic Agentic AI workloads (50-5000 components) plus public vulnerability datasets, yet no analysis is given of how these map to real-world activation conditions or of the resulting false-positive/negative rates for exploitability. This is load-bearing for the claim that the approach yields reliable advisories beyond the synthetic regime.","section":"Evaluation description"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments regarding the evaluation description. We address each major comment below.","responses":[{"response":"We agree that the evaluation section lacks these details. In the revised manuscript we will expand the evaluation section to include the exploitability computation methods (including algorithms and formulas), specific quantitative results and metrics from the ~10000 entries, error handling procedures, and validation outcomes from deterministic replay.","revision_made":"yes","referee_comment":"[Evaluation description] Evaluation description (abstract and evaluation section): The manuscript states that the evaluation uses approximately 10000 component entries but provides no details on computation methods for exploitability, specific results, error handling, or validation outcomes. This absence makes it impossible to assess whether the static-plus-runtime evidence combination supports the central claim of reliable exploitability computation from artefacts, activation conditions, and policies."},{"response":"The evaluation deliberately uses synthetic workloads to support deterministic replay and controlled experimentation. We will add a limitations subsection that discusses the design of the synthetic workloads, their intended approximation to real-world agentic AI activation conditions, and a qualitative assessment of possible false-positive/negative implications. Quantitative false-positive/negative rates from real-world deployments are not available in the current study.","revision_made":"partial","referee_comment":"[Evaluation description] Evaluation description (abstract and evaluation section): The framework is tested exclusively on synthetic Agentic AI workloads (50-5000 components) plus public vulnerability datasets, yet no analysis is given of how these map to real-world activation conditions or of the resulting false-positive/negative rates for exploitability. This is load-bearing for the claim that the approach yields reliable advisories beyond the synthetic regime."}],"tokens_in":1284,"tokens_out":375,"duration_ms":30466,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper presents a framework that binds SBOM and AIBOM artifacts to deterministic environment capture and structured runtime telemetry, then computes exploitability from declared components, observed activation conditions, and enforced policies to produce cryptographically signed CSAF-VEX advisories validated by deterministic replay.\n\nWhat is new is the specific extension of these standards to agentic AI systems, where execution policies and replay validation are meant to make the advisories more precise than static scans alone. The description of how static and runtime evidence combine into signed outputs is clear enough to follow as a protocol outline.\n\nThe evaluation uses roughly 10,000 component entries drawn from synthetic Agentic AI workloads (50 to 5000 components) plus public datasets like OSV, GitHub Advisory, KEV, and EPSS. This setup lets the authors demonstrate the pipeline end-to-end on a controlled scale.\n\nThe main limitation is that everything rests on synthetic workloads. Nothing in the described evaluation shows how the exploitability signals perform against actual deployed agentic systems or measures false-positive and false-negative rates in realistic activation conditions. Without those numbers or comparisons to existing methods, it is hard to judge whether the combined evidence actually improves on current practice.\n\nThis is for readers working on AI supply-chain security and vulnerability management who want to see how existing standards might be adapted to dynamic AI environments. It deserves a serious referee because the protocol idea targets a genuine gap, even though the current evidence base is narrow. I would send it to review so the authors can add concrete validation data or clarify the synthetic-to-real gap.","headline":"Framework for signed CSAF-VEX advisories in agentic AI via AIBOM and runtime telemetry, evaluated only on synthetic workloads of 50-5000 components.","tokens_in":2228,"tokens_out":398,"would_cite":false,"duration_ms":18265,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A protocol-driven framework binds SBOM and AIBOM to deterministic runtime telemetry to compute exploitability and generate signed CSAF-VEX advisories for agentic AI.","keywords":["SBOM","AIBOM","CSAF-VEX","agentic AI","exploitability","runtime telemetry","security advisories","reproducible framework"],"falsifier":"Deploy the framework on a live agentic AI system containing a known exploitable component under controlled conditions and check whether the generated VEX advisory correctly flags or clears the component compared with observed exploit success.","tokens_in":2521,"feed_emoji":"","tokens_out":713,"duration_ms":26080,"temperature":0.7,"pith_summary":"The paper presents a framework that links software and AI bill-of-materials records to captured execution environments and runtime observations. Exploitability is derived from the combination of declared components, activation conditions, and enforced policies, which then feeds into automatically produced, cryptographically signed CSAF-VEX documents. These documents are checked for consistency through deterministic replay. The approach is demonstrated on roughly 10,000 component entries drawn from synthetic agentic AI workloads ranging from 50 to 5,000 components and standard vulnerability datasets. A sympathetic reader would care because it offers a reproducible path from static and dynamic evidence to machine-readable security advisories without relying solely on manual analysis.","feed_headline":"Framework generates signed VEX advisories from runtime telemetry in agentic AI","feed_subtitle":"Binds SBOM and AIBOM artefacts to deterministic capture and execution policies, evaluated on 10,000 components across workloads of 50 to 500","key_machinery":"The AIBOM-driven CSAF-VEX protocol that links declared artefacts and runtime telemetry to signed advisory generation.","core_discovery":"A protocol-driven framework binds SBOM and AIBOM artefacts to deterministic environment capture and structured runtime telemetry. Exploitability is computed from declared artefacts, observed activation conditions, and enforced execution policies. CSAF VEX advisories are generated from the combined static and runtime evidence, cryptographically signed, and validated through deterministic replay.","pith_inferences":["If the synthetic workloads generalize, the same binding of artefacts to telemetry could support continuous advisory updates in deployed production agentic systems.","The approach might reduce reliance on manual triage by surfacing only those vulnerabilities that match observed activation conditions.","Integration with existing SBOM tooling could allow the framework to be inserted into existing CI/CD pipelines for AI components without new data formats."],"forward_implications":["Advisories can be produced automatically from the combination of static artefacts and observed runtime conditions rather than static analysis alone.","Cryptographic signing and deterministic replay make the generated CSAF-VEX documents reproducible across independent verifiers.","Evaluation across workloads of 50 to 5000 components shows the framework scales to moderate-sized synthetic agentic systems while incorporating OSV, GitHub Advisory, KEV, and EPSS data.","Exploitability calculations incorporate enforced execution policies, allowing policy changes to affect advisory output directly."],"fun_headline_variants":["Signed VEX generated from AIBOM and runtime telemetry in agentic AI","AIBOM framework automates CSAF VEX advisories with execution policies","Agentic AI VEX advisories validated via deterministic replay of telemetry","SBOM AIBOM artefacts drive signed VEX for agentic AI component workloads"],"cache_read_input_tokens":64,"weakest_assumption_plain":"Synthetic agentic AI workloads of 50 to 5000 components plus public vulnerability datasets accurately represent real-world execution conditions, and static plus runtime evidence can reliably compute exploitability without significant false positives or negatives.","fun_headline_variants_meta":{"raw":{"variants":["Signed VEX generated from AIBOM and runtime telemetry in agentic AI","AIBOM framework automates CSAF VEX advisories with execution policies","Agentic AI VEX advisories validated via deterministic replay of telemetry","SBOM AIBOM artefacts drive signed VEX for agentic AI component workloads"]},"model":"grok-4.3","cost_usd":0.004314,"raw_usage":{"total_tokens":2101,"prompt_tokens":536,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":43137000,"prompt_tokens_details":{"text_tokens":536,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1492,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":536,"tokens_out":73,"duration_ms":15312,"temperature":1.0,"reasoning_tokens":1492,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T23:43:29.957012+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Deploy the framework on a live agentic AI system containing a known exploitable component under controlled conditions and check whether the generated VEX advisory correctly flags or clears the component compared with observed exploit success.","supporting_citations":[],"review_version":1}