{"id":"88606e32-2ae7-474d-8582-6aed88d63010","arxiv_id":"2606.00145","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"CaB predicts Before/Hit/After boundary tokens to produce auditable switching decisions and boundary-stable control in VLA agents on a Minecraft benchmark under single global calibration.","lead":"The paper proposes Completion at the Boundary (CaB) for vision-language-action agents to decide when instructions are complete using boundary-phase tokens. This targets reliable handoffs in composite tasks under strict deployability constraints with no test-time relearning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Single globally calibrated switching rule may fail to retain usable two-sided evidence across open-ended task polarity shifts","rationale":"The reader’s weakest_assumption directly names the same load-bearing premise. The concrete_test isolates whether the token representation actually mitigates the brittleness the abstract itself acknowledges, moving the verdict from UNVERDICTED to CONDITIONAL pending that check.","tokens_in":1764,"tokens_out":363,"duration_ms":10174,"concrete_test":"On the Minecraft VLA benchmark, partition the test composites into two polarity-matched subsets (one where boundary evidence is predominantly “Before/Hit”, one where it is predominantly “Hit/After”). Apply the identical globally calibrated rule to both subsets and measure whether CaB-When’s F1 on handoff timing drops by >15 % on the second subset relative to the first; if it does, the two-sided evidence claim does not hold under the stated calibration constraint.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that Boundary-Phase Tokens (Before/Hit/After) supply reliable two-sided boundary evidence even when a single switching rule—chosen once on a development set—is frozen and reused on test. The abstract explicitly flags that collapsing asymmetric evidence into a scalar is brittle under polarity shifts; CaB is offered as the fix that preserves two-sided information under the low-calibration discipline. No section or equation in the provided text demonstrates that the token representation remains informative when task polarity (e.g., “do A then stop” vs. “do A then continue”) varies arbitrarily in open-ended instruction spaces. If the tokens’ predictive utility degrades under such shifts, the intervention-aware E1/E2 gains cannot be attributed to CaB rather than to the protocol itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes Completion at the Boundary (CaB) for vision-language-action (VLA) agents to decide instruction completion in short composites under a deployable low-calibration regime (no test-time relearning, single globally calibrated switching rule chosen once on development set). It introduces Boundary-Phase Tokens (Before/Hit/After) to retain two-sided boundary evidence, with CaB-When converting the object to a switching decision and CaB-How reusing it to condition actions for boundary-stable control. Using an intervention-aware E1/E2 protocol, the paper claims improvements in composite execution and handoff quality on a first-person Minecraft VLA benchmark under matched capacity and deployability constraints.","tokens_in":1914,"tokens_out":374,"duration_ms":16664,"significance":"If the results hold, the work addresses an operational gap in deployed VLA systems for open-ended instructions by enabling reliable handoffs without per-task recalibration. The low-calibration discipline and intervention-aware E1/E2 protocol are practical strengths that could support reproducible evaluation in robotics.","major_comments":[{"comment":"Abstract: the central empirical claim of improved composite execution and handoff quality is stated without any quantitative results, baselines, error bars, or metric definitions, so the magnitude and attribution of gains cannot be assessed from the text.","section":"Abstract"},{"comment":"Abstract (and sections describing the low-calibration regime and CaB design): the assertion that Boundary-Phase Tokens retain usable two-sided evidence under arbitrary task polarity shifts with a single frozen switching rule is not supported by any derivation, ablation, or empirical breakdown; this is load-bearing for attributing E1/E2 gains to CaB rather than the protocol itself.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and commit to revisions that improve clarity and attribution without altering the core claims or experimental design.","responses":[{"response":"We agree that the abstract would benefit from explicit quantitative anchors. In the revised version we will insert concise results (e.g., composite success rate deltas, handoff-quality scores, baseline comparisons, and standard-error ranges) together with one-sentence metric definitions, while preserving the abstract length limit.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central empirical claim of improved composite execution and handoff quality is stated without any quantitative results, baselines, error bars, or metric definitions, so the magnitude and attribution of gains cannot be assessed from the text."},{"response":"The intervention-aware E1/E2 protocol already enforces the single frozen rule across the test distribution, and the reported gains are measured under that constraint. Nevertheless, we accept that an explicit polarity-shift ablation would strengthen the causal link to the two-sided Boundary-Phase Tokens. We will add a targeted breakdown (performance stratified by task polarity) and a short derivation sketch of why the three-token representation preserves evidence under sign flips; these additions will appear in the main text and appendix.","revision_made":"yes","referee_comment":"[Abstract] Abstract (and sections describing the low-calibration regime and CaB design): the assertion that Boundary-Phase Tokens retain usable two-sided evidence under arbitrary task polarity shifts with a single frozen switching rule is not supported by any derivation, ablation, or empirical breakdown; this is load-bearing for attributing E1/E2 gains to CaB rather than the protocol itself."}],"tokens_in":1399,"tokens_out":375,"duration_ms":18119,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main move is to represent completion as Boundary-Phase Tokens (Before/Hit/After) instead of collapsing everything to one scalar. That split into CaB-When for the switch decision and CaB-How for boundary-stable action generation is the concrete new piece. It directly targets the closed-loop nature of switching in short composites like \"do A then B.\"\n\nThe low-calibration setup—one rule chosen on dev and reused on test with no test-time learning—is a realistic constraint for open-ended instruction spaces, and the intervention-aware E1/E2 protocol on the Minecraft VLA benchmark fits the deployment focus. The claim that this improves composite execution and handoff quality under matched capacity is the central result.\n\nThe soft spot is the complete absence of quantitative results, baselines, or error bars in the abstract. Without those, it is impossible to tell how large the improvement is or whether it comes from the token representation rather than the protocol itself. The stress-test concern about polarity shifts is worth checking: if the tokens lose predictive value when instructions flip between \"stop after A\" and \"continue after A,\" the two-sided evidence advantage may not hold across the full range of open-ended tasks. The paper would need to demonstrate that the tokens remain informative under such variation.\n\nThis is for robotics researchers working on VLA deployment and chained-instruction reliability. A reader already dealing with handoff failures in similar systems would find the interface design useful.\n\nIt deserves peer review because the problem is practical and the low-calibration discipline is clearly stated, even though the current evidence is too thin to assess the size of the contribution.","headline":"CaB gives a token-based way to keep two-sided completion evidence for VLA handoffs under a single frozen switching rule, but the abstract supplies no numbers so the actual gains stay hard to judge.","tokens_in":2406,"tokens_out":418,"would_cite":false,"duration_ms":18105,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Predicting Before/Hit/After tokens lets VLA agents switch tasks and stabilize control with one fixed calibration rule.","keywords":["vision-language-action","task completion","boundary detection","switching control","composite instructions","Minecraft benchmark","low-calibration deployment"],"falsifier":"On the Minecraft VLA benchmark, run the same E1/E2 protocol with CaB replaced by a scalar completion predictor that uses the identical single global calibration rule; if composite success rate and handoff quality show no gain or a loss, the central claim is false.","tokens_in":2656,"feed_emoji":"🔄","tokens_out":704,"duration_ms":16619,"temperature":0.7,"pith_summary":"Vision-language-action agents often fail at short composite instructions because they cannot reliably detect when one sub-task ends and the next must begin. The paper establishes that collapsing boundary evidence into a single scalar is brittle, so it instead outputs three Boundary-Phase Tokens that keep evidence from both sides of the transition. These tokens feed two modules: one decides the exact switch moment and the other conditions ongoing actions to remain stable across the handoff. All of this is required to hold under the strict deployability rule of a single calibration chosen once on development data and never changed at test time.","feed_headline":"Three-phase tokens improve VLA task switches with one calibration","feed_subtitle":"Boundary-Phase Tokens let agents decide when to hand off and keep actions stable, using only a fixed global rule on the Minecraft benchmark.","key_machinery":"Boundary-Phase Tokens, a three-state prediction (Before/Hit/After) that supplies two-sided boundary evidence for both the switching decision and the action-conditioning step.","core_discovery":"The central claim is that Completion at the Boundary predicts a three-state completion object called Boundary-Phase Tokens (Before/Hit/After) that retains two-sided evidence around each instruction boundary. CaB-When turns this object into an auditable switch decision while CaB-How reuses the same object to condition action generation for boundary-stable control. When evaluated with an intervention-aware E1/E2 protocol on a first-person Minecraft VLA benchmark, the method produces higher composite execution success and better handoff quality than scalar baselines under matched model capacity and the low-calibration constraint of one unchanging global rule.","pith_inferences":["The same three-phase object could be attached to existing VLA models to reduce error accumulation across longer instruction chains.","Retaining asymmetric evidence may help in any sequential decision setting where polarity of observations changes across a boundary.","If the tokens prove stable, downstream planners could treat the Hit state as an explicit synchronization point rather than an implicit switch."],"forward_implications":["Composite task success rises because switches occur at better-timed moments.","Handoff quality improves because action generation is conditioned on the three-phase evidence through the transition.","The same completion object serves both the when decision and the how control without extra modules.","All gains hold under the constraint of no test-time relearning and one fixed calibration rule."],"fun_headline_variants":["CaB predicts Boundary-Phase Tokens for VLA switching","Boundary-Phase Tokens retain two-sided evidence in VLA","Three-phase tokens for low-calibration VLA handoffs","CaB uses tokens for boundary-stable VLA action control"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"A single switching rule chosen once on development data remains effective for open-ended instructions while the three-phase tokens still supply usable evidence on both sides of each boundary.","fun_headline_variants_meta":{"raw":{"variants":["CaB predicts Boundary-Phase Tokens for VLA switching","Boundary-Phase Tokens retain two-sided evidence in VLA","Three-phase tokens for low-calibration VLA handoffs","CaB uses tokens for boundary-stable VLA action control"]},"model":"grok-4.3","cost_usd":0.00693,"raw_usage":{"total_tokens":3170,"prompt_tokens":742,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":69303000,"prompt_tokens_details":{"text_tokens":742,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2363,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":742,"tokens_out":65,"duration_ms":13891,"temperature":1.0,"reasoning_tokens":2363,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T22:36:47.730827+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"On the Minecraft VLA benchmark, run the same E1/E2 protocol with CaB replaced by a scalar completion predictor that uses the identical single global calibration rule; if composite success rate and handoff quality show no gain or a loss, the central claim is false.","supporting_citations":[],"review_version":1}