{"id":"dd74cbd7-411c-4fcb-bd45-382144a9873c","arxiv_id":"2607.27085","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"In 78 controlled mandatory merges, tAVs’ lead and lag gaps converge near lane crossing while projected collision risk peaks at physical lane entry and is dominated by the target-lane leader.","lead":"Controlled public-road trials show transitional automated vehicles converge to similar lead–lag gaps by lane crossing, while projected collision risk peaks at lane entry and is mostly with the target-lane leader. The released-style NC-tALC dataset gives modelers and safety analysts empirical benchmarks for full mandatory lane-change maneuvers, not just gap acceptance.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"SAE risk and preferred-region claims rest on fixed kinematics and same-sample bounds from one supervised stack under forced small gaps.","rationale":"The reader correctly isolates the load-bearing step: turning controlled descriptive trajectories into general tAV “preferred region” and operational collision-risk characterizations via fixed SAE parameters, subjective grouping, and same-sample bound extraction. That is the softest link in the strongest claim; design clarity, timestamp definitions, and the value of a repeatable 78-trial public-road dataset are not in serious doubt. I do not find a deeper internal inconsistency (e.g., the lead-gap formula when X is ahead of A is explicitly flagged; LET/LCC geometry is defined). The concern is external validity and interpretive overreach already flagged in the paper’s limitations—hence CONDITIONAL stays appropriate, not REJECT. Agreement with the reader is full on the weakest assumption; no verdict shift is warranted beyond keeping limitations welded to any general tAV language and requiring sensitivity plus data/code release for stronger confidence.","tokens_in":14365,"tokens_out":713,"duration_ms":14709,"concrete_test":"Recompute SAE_LC (Eq. 4–7) and at-risk fractions (Figs. 8–11) on the same 62 Gap-1/2 trials under a sensitivity grid: τ_e ∈ {0.5,1.0,1.5} s and d ∈ {0.5g,0.6g,0.8g}, and re-extract LCC lead/lag minima after a hold-out or bootstrap split of the three initial-position classes (or leave-one-initial-class-out). If peak-at-LET, front dominance, or post-LCE persistence reverse under milder braking/longer reaction, or if the ~0.54/0.65 s band moves by more than ~0.2 s or loses cross-group consistency, the interpretive claim weakens to sample-specific description.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim packages two process-level patterns—(i) lead/lag convergence by LCC to ~0.54 s / ~0.65 s with lag often larger, and (ii) SAE risk peaking at LET, front-dominated, often persisting past LCE—as characterizations of tAV mandatory LC behavior. Both rest on the Risk Assessment SAE (Eq. 4) with fixed τ_e=0.5 s, d=0.8 g, ℓ_v=5 m, b_0=2 m, plus the GAP ACCEPTANCE section’s post-hoc 0.5 s near-leader/near-follower split and minima read from the same 62 Gap-1/2 trials as a “preferred” operating region. The paper itself notes that longer τ or weaker braking would worsen SAE and that results apply only to the tested system/site/conditions; the experiment also deliberately used relatively small candidate gaps and same-manufacturer tAV followers. If the fixed emergency kinematics mis-rank safety margins, or if the observed band is an artifact of that stack under forced tight merges rather than a controller preference, the behavioral meaning of convergence and front-risk dominance does not transfer beyond this sample—even though the raw trajectory patterns remain useful descriptive facts.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces the NC-tALC dataset from 78 controlled mandatory lane-change trials of a transitional automated vehicle (tAV) on a public road in Apex, NC, using four instrumented vehicles and 20 Hz RTK-GNSS/INS trajectories. It defines key timestamps (ACT, LCS, LET, LCC, LCE and post-LCE offsets), computes lead, lag, and LC time gaps, and applies a spacing-after-emergency (SAE) surrogate to track longitudinal risk. The main empirical claims are that, despite varied initial positions within the target gap, lead and lag gaps evolve toward a relatively narrow band by lane-change crossing (reported minima near 0.54 s lead and 0.65 s lag, with lag often larger thereafter), and that projected collision risk often rises through the maneuver, peaks near left-edge touching, is predominantly front risk versus the target-lane leader, and can persist past lane-change end. The authors position the work as one of the first controlled characterizations of the full mandatory LC process for tAVs that decide and execute the maneuver under supervision.","tokens_in":14624,"tokens_out":1513,"duration_ms":39214,"significance":"If the reported process-level patterns hold under the stated scope, the paper fills a genuine empirical gap: systematic public-road data on tAVs that independently decide and execute mandatory lane changes, rather than Level-4 open datasets, pooled Level 1–2 assisted events, or unreproduced assisted-LC campaigns. The controlled four-vehicle design, high-resolution trajectories, explicit timestamp definitions, and public dataset framing are concrete strengths for model calibration, simulation validation, and safety-method benchmarking. The emphasis on evolution through the full maneuver (not only gap acceptance) is useful for traffic and safety modeling. Credit is due for transparent experimental control of initial spacing classes, clear instrumentation, and explicit scope limitations in the conclusions. The contribution is primarily observational and dataset-centered rather than a new theory; its lasting value depends on careful interpretation of the preferred-region and SAE-risk claims and on dataset usability.","major_comments":[{"comment":"GAP ACCEPTANCE AND EVOLUTION: The conjecture of a controller “preferred lead–lag operating region” at LCC rests on minima (~0.54 s lead, ~0.65 s lag) read from the same near-leader/near-follower trials that define the groups, then applied back as shared bounds that the near-center group also meets. That is a same-sample descriptive pattern, not independent evidence of a preferred setpoint. Please either (i) reframe strictly as an empirical convergence description without “preferred region” language, or (ii) support the claim with out-of-sample checks (e.g., hold-out trials, Gap-0 vs Gap-1/2, or pre-registered bounds) and report group sizes n for near-leader / near-center / near-follower at ACT and at LCC.","section":"GAP ACCEPTANCE AND EVOLUTION"},{"comment":"RISK ASSESSMENT, Eq. (4)–(7): Front-risk dominance and at-risk fractions (e.g., 83.9% SAE_LC<0 at LET; ~1/3 still at risk at LCE) are load-bearing for the safety narrative, yet SAE uses a single fixed emergency parameter set (τ_e=0.5 s, d=0.8 g, ℓ_v=5 m, b_0=2 m) with no sensitivity table. The text correctly notes that longer τ or weaker braking would worsen SAE, but the quantitative claims (peak at LET, persistence past LCE, front dominance percentages) could shift under plausible alternatives. Please add a compact sensitivity analysis (at least τ_e and d) showing whether peak timing, front/rear dominance, and post-LCE persistence are robust, and keep the interpretation as a kinematic surrogate under stated assumptions rather than operational collision probability.","section":"RISK ASSESSMENT"},{"comment":"CONCLUSIONS AND EXPERIMENTS: The abstract and novelty statements characterize “tAV” mandatory LC behavior, while the body correctly limits findings to one supervised stack, one site, deliberately small candidate gaps, and same-manufacturer tAV followers. That scope mismatch is load-bearing for transferability. Please align title/abstract/practical-applications wording with the single-system, controlled-challenge design (e.g., “a commercial tAV under supervised mandatory merge”), and state more clearly that Gap-3 was never selected and that small-gap forcing may inflate risk relative to naturalistic merges.","section":"CONCLUSIONS AND DISCUSSIONS"}],"minor_comments":[{"comment":"Submission date on the title page reads “July 29, 2026” and the arXiv stamp is 29 Jul 2026; confirm this is intentional and consistent with the journal’s dating practice.","section":"Title page"},{"comment":"Eq. (2)–(3): clarify sign convention and when negative lead/lag gaps are retained versus undefined; a short note that vehicle-center headway with fixed ℓ_v can bias short gaps would help reproducibility.","section":"Time gap measurement"},{"comment":"Figures 5–12: axis units, sample sizes per panel/group, and the 0.5 s grouping threshold should appear in captions; several figures are hard to interpret from the text alone.","section":"Figures 5–12"},{"comment":"Table 1 is clear on gap selection, but the main text never tabulates how the 78 trials map into the three ACT lead–lag groups used in §GAP ACCEPTANCE; add a small contingency table.","section":"TABLE 1 / GAP ACCEPTANCE"},{"comment":"Minor prose issues: spacing typos (“throughoutthe”, “lane-changeprocess”), inconsistent “tAV dataset” vs “NC-tALC”, and mixed “lead–lag” hyphenation. A copy-edit pass would help.","section":null},{"comment":"Related work: the JRC/Mattas campaign and TGSIM are well cited; if the NC-tALC companion arXiv (Sharma et al. 2026) is the data release, state access conditions and what is public versus reserved for the ACC-follower paper.","section":"INTRODUCTION"}],"recommendation":"minor_revision","confidential_remarks":"The empirical core and dataset appear publishable with tightened interpretation; I do not see a fatal methodological error. The main risk is over-general branding of one supervised OEM stack under forced tight merges as “tAV lane-changing behavior.” If the journal prioritizes reusable datasets and process-level AV behavior, minor revision is appropriate. I did not treat the skeptic’s SAE/preferred-region concern as grounds for reject: the paper already caveats scope, and the raw trajectory patterns remain useful if claims are narrowed and sensitivity is added."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is simple: they ran a real public-road, four-vehicle mandatory merge experiment (n=78) with RTK/INS at 20 Hz, forced different initial positions in small target gaps, and tracked the full maneuver—not just gap acceptance. That fills a hole relative to Waymo L4 snippets, TGSIM assisted events, and the JRC assisted-LC campaign (few system-initiated cases, data not public).\n\nWhat they do well is the craft. Activation, LCS/LET/LCC/LCE definitions are explicit; lead/lag/LC gaps are cleanly defined; the descriptive pattern is credible on the sample: different starts tighten toward a common lead–lag band by centerline crossing, lag often larger than lead after that, and projected longitudinal risk (their SAE) often worst at first physical entry, mostly vs the target leader, sometimes still negative after LCE. They also say out loud that results are one supervised system, one site, challenging gaps, and that softer braking or slower reaction would look worse. That honesty helps.\n\nSoft spots are real but proportionate. The ~0.54 s / ~0.65 s bounds and “preferred operating region” are same-sample minima plus a subjective 0.5 s grouping, then read as controller intent—mild circularity, not fraud. SAE with τ=0.5 s and 0.8 g is a fixed surrogate, not observed crash risk; front-risk dominance is interesting under those assumptions, not a general tAV law. “tAV” is a useful label here, not a settled SAE class. Code/raw release is not evidenced in-text, which matters for a dataset paper. None of that kills the trajectories or the process-level plots.\n\nWho it’s for: people building or validating mandatory LC models, AV safety metrics, and mixed-traffic sims who need full-maneuver benchmarks rather than acceptance-instant snapshots. Math is elementary kinematics; citations look fair to the prior AV LC work. I’d send it to referees. Engage if you work this area; treat the numbers as this stack under these conditions, not universal tAV behavior.","headline":"Useful controlled mandatory-merge trajectories and process-level patterns; the “preferred region” and SAE risk story should stay tied to one stack and fixed kinematics.","tokens_in":15379,"tokens_out":538,"would_cite":true,"duration_ms":18702,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"In controlled public-road trials, transitional automated vehicles converge lead and lag gaps by lane crossing while front collision risk peaks at physical lane entry and can outlast the maneuver.","keywords":["transitional autonomous vehicles","mandatory lane changing","lead-lag gaps","surrogate safety measures","controlled field experiment","trajectory dataset","collision risk evolution"],"falsifier":"Repeat the same mandatory-merge protocol with other manufacturers or software versions, or with weaker assumed braking and longer reaction times: if lead–lag states no longer converge near those bounds by crossing, or if rear risk dominates or risk vanishes by maneuver end, the claimed convergence and front-risk pattern fail.","tokens_in":15107,"feed_emoji":"🚗","tokens_out":963,"duration_ms":19159,"temperature":0.7,"pith_summary":"This paper introduces a controlled public-road dataset of 78 mandatory lane changes by transitional automated vehicles—systems that decide and execute lane changes under driver supervision—and uses the trajectories to map behavior across the full maneuver, not only at gap acceptance. Four instrumented vehicles created repeatable initial spacings while the lane changer picked among candidate target gaps. Despite very different starting lead–lag positions, lead and lag time gaps consistently moved toward a relatively narrow band by the moment the vehicle center crossed into the target lane, with the lag gap often larger than the lead gap thereafter. Projected longitudinal collision risk, measured by an emergency-braking spacing surrogate, rose through the maneuver, peaked when the vehicle’s left edge first entered the target lane, was dominated by conflict with the target-lane leader rather than the follower, and sometimes remained after the lateral maneuver had ended. The practical stake is clear: models, simulators, and safety checks that treat automated lane change as a single acceptance instant will miss both the convergence behavior and the risk that peaks at entry and can persist afterward.","feed_headline":"AV lane changes converge gaps; risk peaks at lane entry","feed_subtitle":"Controlled public-road trials show front risk can outlast the maneuver itself.","key_machinery":"The NC-tALC controlled four-vehicle mandatory-merge experiment with high-rate RTK-GNSS/INS trajectories, annotated key timestamps (activation, lane-change start, left-edge touching, crossing, end, and post-end offsets), lead/lag/LC time gaps, and the spacing-after-emergency (SAE) surrogate that scores front and rear longitudinal risk under a fixed emergency-braking scenario.","core_discovery":"Despite substantially different initial lead–lag conditions, the tested transitional automated vehicles’ lead and lag gaps evolve toward a relatively narrow operating region by lane-change crossing (empirical minima near about 0.54 s lead and 0.65 s lag, with lag often larger than lead thereafter). Significant projected longitudinal collision risk develops during the maneuver, typically peaks at left-edge touching (physical lane entry), is predominantly front risk versus the target-lane leader, and can persist beyond lane-change end.","pith_inferences":["If front-risk tolerance is a controller preference rather than a site artifact, mixed traffic may see more leader braking or cut-in friction than human-centric merge models predict.","Regulators and OEM test suites that only check gap size at initiation or completion would miss the highest-risk window identified here.","Comparing the same protocol under ACC-only followers versus automated followers (reserved in the paper’s companion experiment) would isolate how follower automation changes the risk timeline."],"forward_implications":["Lane-change models for transitional automation should track longitudinal adjustment through the full process, not only gap acceptance or a single timestamp.","Safety evaluation should score front and rear risk separately from physical lane entry through post-maneuver stabilization, because risk can peak at entry and outlast completion.","The dataset supplies empirical benchmarks for calibrating behavioral models and validating simulation of mandatory merges under controlled initial conditions.","Lane-crossing is a behavioral milestone: initially different lead–lag conditions have largely converged and rear risk is largely reduced by that point."],"fun_headline_variants":["tAV lane changes: gaps converge narrow near crossing","Mandatory AV lane risk peaks at physical lane entry","Lead-lag gaps tighten; front collision risk outlasts end","NC-tALC: tAV gaps converge, risk highest at lane touch","Controlled trials: AV lane risk dominated by target leader"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim that these patterns show a preferred operating region and real collision risk rests on one supervised vehicle type, one site, challenging small gaps, a subjective initial-condition grouping, and a fixed emergency-braking formula whose parameters may understate or mislocate risk.","fun_headline_variants_meta":{"raw":{"variants":["tAV lane changes: gaps converge narrow near crossing","Mandatory AV lane risk peaks at physical lane entry","Lead-lag gaps tighten; front collision risk outlasts end","NC-tALC: tAV gaps converge, risk highest at lane touch","Controlled trials: AV lane risk dominated by target leader"]},"model":"grok-4.5","effort":"low","cost_usd":0.002399,"raw_usage":{"total_tokens":1043,"prompt_tokens":856,"num_sources_used":0,"completion_tokens":66,"cost_in_usd_ticks":23988000,"prompt_tokens_details":{"text_tokens":856,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":121,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":856,"tokens_out":66,"duration_ms":3840,"temperature":1.0,"reasoning_tokens":121,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T11:35:17.814752+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the same mandatory-merge protocol with other manufacturers or software versions, or with weaker assumed braking and longer reaction times: if lead–lag states no longer converge near those bounds by crossing, or if rear risk dominates or risk vanishes by maneuver end, the claimed convergence and front-risk pattern fail.","supporting_citations":[],"review_version":1}