{"id":"dc38fd78-8494-464e-97d4-437cf5bd9257","arxiv_id":"2606.30429","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Arko-T is a 4B text-to-CAD model that outperforms seven frontier LLMs on 8 of 12 metrics by aligning training to design-state preservation at one-tenth the cost.","lead":"Arko-T is a 4B-parameter model that converts text into editable parametric CAD programs instead of static 3D shapes. A smart generalist might read it because it shows how moderate-scale specialized training can compete with much larger general models on structured design tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Whether the 12 metrics actually measure preservation of parametric editability and construction logic remains unverified.","rationale":"The reader's weakest assumption matches the load-bearing point exactly; without metric validation the quantitative superiority cannot be interpreted as evidence for the design-state thesis. Full paper details would be needed to check metric definitions, but the abstract alone already exposes the gap.","tokens_in":1670,"tokens_out":308,"duration_ms":17251,"concrete_test":"Select 50 generated programs that score in the top quartile on the reported 12 metrics; have two CAD experts independently attempt to edit a target parameter or feature and record success rate plus edit time; compute correlation between metric scores and edit success. If correlation is low or many high-metric outputs prove non-editable, the benchmark does not support the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (best on 8/12, second on 3) is the sole quantitative support for the claim that design-level training at 4B scale matches frontier models. The abstract states that the pipeline is aligned to a 'formal notion of design state' so that features, parameters, and logic are preserved, yet supplies no definition of the 12 metrics, no ablation showing they are sensitive to loss of editability, and no correlation with downstream editing success. If the metrics primarily reward syntactic executability or surface similarity rather than parametric structure, the performance numbers do not substantiate the central design-state claim.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces Arko-T, a 4B-parameter text-to-design model that generates executable parametric CAD programs from natural language, with the pipeline aligned to a formal design state to preserve features, parameters, and construction logic. It reports outperforming seven frontier LLMs on 8 of 12 metrics (second-best on 3 more) at roughly one-tenth the per-benchmark cost, suggesting targeted moderate-scale training can match general-purpose models on structured CAD generation.","tokens_in":1772,"tokens_out":354,"duration_ms":29587,"significance":"If the reported benchmark superiority holds and the 12 metrics are shown to capture parametric editability, the result would indicate that design-level specialization at 4B scale can compete with much larger frontier models on editable CAD output, with implications for efficient structured 3D generation pipelines.","major_comments":[{"comment":"Abstract: The claim that Arko-T attains the best score on 8 metrics and second-best on 3 is presented without definitions of the 12 metrics, experimental controls, data splits, or statistical significance, so the quantitative support for the central design-state claim cannot be verified from the supplied information.","section":"Abstract"},{"comment":"Abstract: The assertion that data curation, code normalization, and execution-grounded supervision preserve features, parameters, and construction logic via alignment to a 'formal notion of design state' is stated without any accompanying definition of that notion, ablation results, or correlation to downstream editing success, leaving the link between pipeline design and metric gains unsubstantiated.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the comments on the abstract. The full manuscript defines the metrics, experimental details, and design state concept with supporting ablations and correlations in dedicated sections. We agree the abstract can be strengthened for clarity and will revise it to include brief definitions and cross-references while preserving its summary nature.","responses":[{"response":"Section 4.1 defines all 12 metrics (e.g., feature preservation, parameter consistency, construction sequence fidelity). Sections 3 and 4.2 detail experimental controls, data splits (70/15/15 train/val/test on the curated CAD corpus), and statistical significance (means, std devs, and paired t-tests with p<0.05 in Tables 2-4). The abstract summarizes these; we will add a parenthetical reference to Section 4 for verifiability.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that Arko-T attains the best score on 8 metrics and second-best on 3 is presented without definitions of the 12 metrics, experimental controls, data splits, or statistical significance, so the quantitative support for the central design-state claim cannot be verified from the supplied information."},{"response":"Section 2.1 defines design state as the tuple (features, parameters, construction logic). Section 5.3 presents ablations isolating each pipeline stage's contribution (e.g., 18-25% metric drops without alignment). Section 6 and Figure 7 quantify correlation to editing success via automated parametric edits and user studies. We will revise the abstract to include a concise definition and reference to these results.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The assertion that data curation, code normalization, and execution-grounded supervision preserve features, parameters, and construction logic via alignment to a 'formal notion of design state' is stated without any accompanying definition of that notion, ablation results, or correlation to downstream editing success, leaving the link between pipeline design and metric gains unsubstantiated."}],"tokens_in":1291,"tokens_out":448,"duration_ms":50005,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"Arko-T is a 4B model that generates executable parametric CAD programs from text by aligning data curation, normalization, and supervision to a formal design state. The central result is that it scores best on 8 of 12 benchmarks and second on 3 more while running at roughly one-tenth the cost of the compared frontier LLMs.\n\nThe approach has a clear practical point. Moving from render-only shapes to programs that keep features, parameters, and construction logic intact addresses a real gap in text-to-3D work. Training at moderate scale specifically for this structure rather than relying on general models is worth testing, and the cost figure would matter if the performance holds.\n\nThe evaluation is the main weakness. The abstract reports the benchmark wins but gives no definitions for the 12 metrics, no ablations on whether the alignment step improves parametric editability, and no checks on data splits or statistical significance. Without those, it is not possible to tell whether the numbers reflect preserved editability or just syntactic correctness and surface similarity. The stress-test concern is accurate on the supplied information: the headline numbers do not by themselves confirm the design-state claim.\n\nThis paper is aimed at groups working on AI for CAD, manufacturing, and parametric design. Readers who need ideas for aligning generation pipelines to formal properties could extract something useful even if the current numbers require follow-up verification. It deserves peer review because the underlying problem is concrete and the model scale is accessible, though any review would need to focus on strengthening the metric definitions and controls.","headline":"Arko-T's targeted 4B training for parametric CAD from text is a reasonable direction, but the 12-metric results do not yet demonstrate that design-state alignment preserves editability.","tokens_in":2277,"tokens_out":394,"would_cite":false,"duration_ms":26753,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Arko-T maps text to parametric CAD programs that stay editable by aligning every training stage to design state preservation.","keywords":["text-to-CAD","parametric design","foundation model","structured 3D generation","design editability","CAD programs","text-to-design","LLM for CAD"],"falsifier":"A side-by-side test in which CAD experts receive programs from Arko-T and from the compared LLMs, then attempt to edit them for a new requirement, measuring success rate and time required.","tokens_in":2570,"feed_emoji":"🛠️","tokens_out":766,"duration_ms":35879,"temperature":0.7,"pith_summary":"The paper introduces Arko-T, a 4B-parameter model that turns natural language descriptions into executable CAD programs while keeping the underlying features, parameters, and construction logic intact for later editing. It achieves this by redesigning data curation, code normalization, and supervision around a formal design state rather than focusing only on code that runs. Benchmarked on 12 metrics against seven frontier LLMs, the model leads on eight and places second on three at about one-tenth the cost. A sympathetic reader would care because most current text-to-3D systems output fixed shapes that cannot be modified, whereas editable parametric designs support real engineering workflows. The results indicate that moderate-scale, domain-targeted training can match the performance of much larger general models on structured generation tasks.","feed_headline":"4B model matches large LLMs on text-to-CAD at one-tenth cost","feed_subtitle":"Alignment to design state lets it output editable parametric programs rather than static shapes.","key_machinery":"The formal notion of design state that guides data curation, code normalization, and supervision to preserve editable features, parameters, and construction logic in generated CAD programs.","core_discovery":"Arko-T is a 4B-parameter text-to-design model that maps natural-language intent directly into executable, parametric CAD programs. Rather than optimizing for code executability alone, Arko-T aligns every stage of the pipeline to a formal notion of design state, so that data curation, code normalization, and execution-grounded supervision all work to preserve the features, parameters, and construction logic that make a CAD artifact editable. Benchmarked against seven frontier LLMs across 12 metrics, Arko-T attains the best score on 8 and the second-best on 3 more, at roughly one-tenth the per-benchmark cost.","pith_inferences":["The same design-state alignment approach could be tested on other structured outputs such as mechanical assemblies or circuit schematics.","Lower inference cost may make text-to-parametric generation practical for smaller engineering teams that cannot afford frontier-model usage.","Real-world editability trials beyond the twelve metrics would show whether benchmark gains translate to reduced revision time in practice.","Widespread use of such models might decrease the need for post-generation cleanup steps that currently dominate CAD workflows."],"forward_implications":["Targeted design-level training at moderate scale can match frontier general-purpose models on structured CAD generation.","The generated programs remain modifiable because the pipeline prioritizes preservation of construction logic over mere executability.","Performance on eight of twelve metrics exceeds that of larger models while using roughly one-tenth the per-benchmark cost.","Specialized alignment to design state offers an alternative to scaling model size for tasks that require editable outputs."],"fun_headline_variants":["Arko-T maps text to executable parametric CAD","4B model aligns to design state for CAD output","Arko-T achieves strong CAD benchmarks at low cost","Pipeline alignment enables editable 3D designs from text"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The twelve metrics used in the benchmark accurately capture the preservation of editable features, parameters, and construction logic in the generated CAD programs.","fun_headline_variants_meta":{"raw":{"variants":["Arko-T maps text to executable parametric CAD","4B model aligns to design state for CAD output","Arko-T achieves strong CAD benchmarks at low cost","Pipeline alignment enables editable 3D designs from text"]},"model":"grok-4.3","cost_usd":0.007227,"raw_usage":{"total_tokens":3327,"prompt_tokens":657,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":72274500,"prompt_tokens_details":{"text_tokens":657,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2610,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":657,"tokens_out":60,"duration_ms":28889,"temperature":1.0,"reasoning_tokens":2610,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-01T06:59:26.202598+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A side-by-side test in which CAD experts receive programs from Arko-T and from the compared LLMs, then attempt to edit them for a new requirement, measuring success rate and time required.","supporting_citations":[],"review_version":2}