{"id":"e1635e00-b115-4512-8d26-ba0c29388e6e","arxiv_id":"2603.02156","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"On 6G-Bench, deterministic reasoning accuracy rises from 22% at 135M to 71% at 7B, but the best accuracy-per-edge-resource balance occurs at 1.5–3B parameters.","lead":"An empirical study tests how well tiny and small language models (135M–7B parameters) can reason about 6G network decisions, using a 30-task benchmark. It finds accuracy grows unevenly, with a sweet spot around 1.5–3B parameters when latency and memory are counted.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 1–1.5B stability transition is a model-family artifact: Granite-1B already reaches A1=0.559 and Δ5=0.217 at the same nominal size as Llama-1B (A1=0.373, Δ5=0.356), so parameter count does not identify the transition.","rationale":"The reader's CONDITIONAL verdict is appropriate; my stress-test supplies a distinct reason for it. The strongest claim includes a precise parameter threshold, and that threshold is the load-bearing part: the deployment guidance hinges on it. The Δ5 metric mismatch (pass@k on 7 tasks vs A1 on all 3,722 MCQs) and the implausible Granite VRAM/latency values are real and reinforce the need for revision, but the family confound is prior: even if Δ5 were recomputed on a common subset and profiling corrected, the 1B→1.5B comparison would still compare Llama to Qwen. A controlled same-family sweep is the minimal check that can settle whether a scale effect exists. I therefore leave the reader's CONDITIONAL verdict unchanged.","tokens_in":19428,"tokens_out":15639,"duration_ms":128572,"concrete_test":"Run a same-family crossover on the public 6G-Bench repo: add Qwen2.5-0.5B-Instruct (same family as the 1.5/3/7B points) and Llama-3.2-3B, under the identical zero-shot protocol, then recompute A1 and Δ5. If Qwen-0.5B already attains A1≈0.53/Δ5≈0.14, or if the within-1B spread from Table IV (A1:0.186, Δ5:0.139) exceeds or equals the reported 1B→1.5B effect (A1:0.158, Δ5:0.218), the transition is a model-selection artifact and the 1.5–3B deployment tier should be re-derived.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table IV is the entire basis for the scale-threshold conclusion, but it is an unpaired model comparison: N is confounded with pretraining data, alignment, and architecture. At the same nominal N≈1B, Llama-3.2-1B has A1=0.373/Δ5=0.356 while Granite-1B has A1=0.559/Δ5=0.217. The abstract's '1–1.5B stability transition' (A1 +0.158, Δ5 −0.218) is just the difference between Llama-1B and Qwen-1.5B; the within-1B spread is A1 +0.186 and Δ5 −0.139, i.e. comparable to or larger than the claimed cross-scale jump. Taking the best model at each size, A1 actually decreases from 0.559 (Granite-1B) to 0.531 (Qwen-1.5B). The paper's own residual analysis admits architecture-dependent deviations of ±0.1, so the fitted log-linear curve cannot separate scale from model identity. Therefore the '1.5–3B stability tier' does not follow: Granite-1B (Δ5=0.217) and LFM2.5-1.2B (Δ5=0.146) are already as stable as Qwen-1.5B (Δ5=0.138) at lower resource cost. This does not question the reported accuracies; it questions the inference from unpaired points to a parameter threshold.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical scaling study of ten instruction-tuned small language models (135M–7B parameters) on 6G-Bench, the authors' 30-task / 3,722-MCQ benchmark for network-level semantic reasoning in AI-native 6G. It measures deterministic accuracy pass@1, stochastic pass@3/pass@5, and an instability gap Δ5 = A5 − A1; fits a log-linear accuracy-vs-parameter curve; defines group-level scaling sensitivity; profiles single-query latency and peak VRAM; and introduces an Edge Score ES = A1/(L·M). The headline claims are that deterministic accuracy rises from 0.224 (SmolLM2-135M) to 0.707 (Qwen2.5-7B), that a stability transition occurs in the 1–1.5B range, and that mid-scale 1.5–3B models achieve the most favorable balance between deterministic stability and edge efficiency.","tokens_in":19870,"tokens_out":5698,"duration_ms":54113,"significance":"If the empirical regularities hold, the deployment-tier guidance would be a useful input to 6G architecture discussions: ultra-compact models for localized decisions, 1.5–3B for stability-critical edge reasoning, and 7B for centralized orchestration. The paper's strengths are empirical: a reproducible public benchmark, local evaluation of all checkpoints under a unified prompting/extraction protocol, explicit binomial CIs for pass@1, and an honest limitations section. It does not propose new theory; the scaling fit is descriptive. The main risks are that the headline '1–1.5B stability threshold' conflates model identity with parameter count, and that the Edge Score ranking is sensitive to profiling artifacts. Because the underlying data and code are public, these issues are fixable by reframing the claims and adding the missing uncertainty/sensitivity analysis; the study remains a potentially valuable empirical contribution.","major_comments":[{"comment":"Table III states that pass@k is evaluated only on 7 tasks (T2, T9, T12, T19, T20, T26, T30), while Table IV reports Δ5 = A5 − A1 with 95% CIs computed by variance summation of A1 and A5, as if both terms cover the full M=3,722 questions. Eq. (3) defines A_k over the whole task set T. Thus A5 and A1 are not commensurate: the gap can shrink merely because the 7-task subset is easier under greedy decoding. The paper never states this restriction in Table IV or in the abstract. This invalidates the reported contraction of Δ5 from 0.365 to 0.031 as a measure of reasoning instability and undermines the stability-transition conclusion.","section":"§III-D, Eq. (3); Table III vs Table IV"},{"comment":"The '1–1.5B stability transition' is inferred from a single pair of checkpoints, Llama-3.2-1B versus Qwen2.5-1.5B. Parameter count is fully confounded with pretraining data, alignment, and architecture. Within the same nominal N≈1B, Granite-1B reaches A1=0.559 and Δ5=0.217 while Llama-1B reaches A1=0.373 and Δ5=0.356; the within-1B spread is comparable to or larger than the claimed cross-scale jump, and taking the best model at each size, A1 decreases from Granite-1B (0.559) to Qwen-1.5B (0.531). The paper's own residuals show architecture-dependent deviations of ±0.1. A z-test between two specific checkpoints is not evidence of a parameter threshold. The central 'mid-scale stability tier' therefore does not follow from Table IV.","section":"§IV-A, Table IV; abstract"},{"comment":"Edge Score = A1/(L·M) is only as credible as the latency and memory profiles. Table VI reports Granite-350M with 16.194 GB peak VRAM and 1814.7±128.7 ms latency, and Granite-1B with 20.371 GB and 3161.3±63.3 ms — values an order of magnitude worse than the 7B models (≈1 GB, ≈340–390 ms). This is not a plausible scaling of computational cost and suggests implementation-dependent artifacts in the inference harness. Since ES is a ratio, Figure 7's ranking is dominated by these artifacts. No uncertainty is propagated to ES, and the deployment-tier recommendations in §IV-D rest on this unvalidated ranking.","section":"§III-D.3, Eq. (8); Table VI; Fig. 7"}],"minor_comments":[{"comment":"Table III lists 11 models, but Table IV and the abstract evaluate 10; LFM2-8B appears only in Table VI. Please clarify the model count and the role of LFM2-8B.","section":"Table III vs Table IV"},{"comment":"The caption should explicitly state that A3, A5, and Δ5 are computed on the 7-task subset; currently the notation implies the full 3,722 MCQs.","section":"Table IV caption"},{"comment":"Eq. (4) defines A(N)=α log(1+N)+β, but Table IV's residual line says A1−(α log N+β). Align the two formulations.","section":"Eq. (4) vs Table IV residual"},{"comment":"Add error bars or a note that Edge Score does not include uncertainty; currently the bars imply precision that the underlying measurements do not support.","section":"Fig. 7"},{"comment":"The limitations paragraph is welcome, but it is not connected back to the headline claims. Add a sentence stating which conclusions are robust to model-family effects and which are not.","section":"Section IV-E"}],"recommendation":"major_revision","confidential_remarks":"The benchmark 6G-Bench is the authors' own work, and the paper does not discuss potential benchmark-author bias or benchmark-overfitting by design. The editor may want to ensure independent scrutiny of benchmark construction and public availability of the exact MCQs before the scaling claims are accepted. The paper would also be strengthened by reproducing the inference profiling on a more homogeneous, controlled stack, since the current Edge Score ranking is dominated by likely artifacts. I do not recommend rejection because the measurement corpus and code are public and the central scaling trend is plausibly real; however, the threshold and efficiency conclusions need substantial rework."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Karin,\n\nYou should know about this one if you care about small-model scaling for edge systems. The paper evaluates ten instruction-tuned models, 135M to 7B, on 6G-Bench (a benchmark from the same authors) and reports deterministic pass@1, stochastic pass@k, and an \"instability gap\" Δ5 = A5 − A1. The empirical pass@1 table is useful and probably reproducible — accuracy does climb from 0.224 to 0.707, and the difference between Llama-1B and Qwen-1.5B is statistically clear. I'd trust those numbers as measurements of these particular models.\n\nBut the interpretive claims do not hold up. The headline \"1–1.5B stability transition\" is an unpaired model comparison: Granite-1B reaches A1=0.559 and Δ5=0.217, while Llama-1B sits at 0.373/0.356. The cross-scale jump is just the difference between two different model families. There is no evidence that parameter count per se causes the transition.\n\nWorse, the Δ5 metric is computed inconsistently. Pass@k is only evaluated on 7 of the 30 tasks (stated in Table III), but Table IV reports Δ5 as if both A5 and A1 came from the full 3,722 MCQs. The contraction from 0.365 to 0.031 could be an artifact of the mixed question sets.\n\nThe efficiency analysis is also shaky. Table VI lists 11 models though the study covers 10, and the Granite-350M/1B entries show 16–20 GB VRAM and 1.8–3.2 s latency — absurd for those sizes, and clearly the reason the Edge Score ranking (Fig. 7) puts them on top. The resid column in Table IV doesn't match the stated log-linear fit α=0.115, β=0.478: for SmolLM2-135M the residual should be about −0.269, not −0.026. That suggests a calculation error in the scaling analysis.\n\nNone of this undermines the raw A1 measurements. But the paper's contribution is the scaling interpretation, and that's where the problems concentrate. The right fix is to recompute Δ5 on matched tasks, analyze model-family confounds (e.g., best-vs-best per size), correct the profiling table, and rerun the residuals. I'd send it to peer review because the topic and the public data are worth referees' time, but I'd expect a major revision. Not a desk-reject.","headline":"Useful raw pass@1 measurements, but the stability threshold and edge-efficiency conclusions are confounded by model-family comparisons and metric inconsistencies.","tokens_in":20333,"tokens_out":7949,"would_cite":false,"duration_ms":61991,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For AI-native 6G networks, semantic reasoning reliability scales non-uniformly: deterministic accuracy jumps sharply between 1B and 1.5B parameters and then plateaus, making 1.5–3B models the best edge-deployment balance.","keywords":["6G","small language models","semantic reasoning","scaling laws","edge deployment","instability gap","AI-native networks","reasoning stability"],"falsifier":"Recompute A1 and A5 on the same 7-task subset used for pass@k and on the full 30-task set; if the 1B-to-1.5B Δ5 collapse disappears on the full set, the stability transition is a subset artifact. Separately, rerun the Edge Score with memory footprint measured at matched sequence lengths and with quantization enabled, and check whether mid-scale models still dominate the frontier.","tokens_in":19337,"feed_emoji":"📡","tokens_out":5436,"duration_ms":51260,"temperature":0.7,"pith_summary":"The paper tries to establish that the right size for an AI reasoning layer in a 6G network is not the largest model that fits, but the smallest model that reasons stably. Evaluating ten instruction-tuned language models from 135 million to 7 billion parameters on 6G-Bench's 30 decision-making tasks, it finds deterministic accuracy rises from 0.224 to 0.707, but the gains are concentrated between 1B and 1.5B, where accuracy jumps from 0.373 to 0.531 and the instability gap collapses. Beyond 3B, improvements shrink to +0.064. Using an Edge Score that divides accuracy by latency and memory, the paper argues that mid-scale models around 1.5–3B give the best reliability per unit edge resource, while tiny models serve only low-risk localized decisions.","feed_headline":"1.5B models hit the 6G reasoning sweet spot","feed_subtitle":"Scaling 135M to 7B shows reliability jumps at 1–1.5B, then diminishing returns; mid-size wins on edge cost.","key_machinery":"The central objects are (1) 6G-Bench, a standardization-aligned benchmark of 30 tasks and 3,722 multiple-choice questions spanning intent and policy reasoning, network slicing, trust and security, agentic control, and distributed intelligence; (2) three metrics: pass@1 for deterministic accuracy, pass@k for stochastic robustness, and the instability gap Δk = Ak − A1; and (3) the Edge Score ES = A1/(L·M), which normalizes accuracy by inference latency and peak VRAM. The argument runs by comparing these quantities across ten models from 135M to 7B parameters and fitting a log-linear scaling curve A(N) = α log(1+N) + β, whose residuals and domain-level sensitivity coefficients reveal that scali","core_discovery":"On the paper's own terms: there is an empirical scaling regime for network-level semantic reasoning in compact language models. Deterministic single-shot accuracy grows monotonically from 0.224 at 135M parameters to 0.707 at 7B, but the growth is not smooth; the 1B-to-1.5B step produces the largest gain, and the instability gap Δ5, defined as pass@5 minus pass@1, contracts from roughly 0.36 in the sub-1B regime to 0.031 at 7B. This contraction is interpreted as convergence between stochastic exploration and deterministic inference, a prerequisite for safety-critical control. Additionally, semantic reliability per unit edge resource, as measured by the Edge Score, does not scale with paramete","pith_inferences":["The Δ5 contraction may be overstated because pass@5 is measured on only 7 of the 30 tasks while pass@1 is averaged over all 3,722 questions; recomputing both metrics on the same task set could move or erase the reported 1B–1.5B threshold.","The Edge Score ranking is a snapshot of bf16, single-GPU, zero-shot inference; under quantization, batched serving, or lower-precision edge hardware, the memory and latency profiles could change which model class wins.","If the non-uniform scaling curve generalizes, the search for 'minimal stable capacity' could become a standard design step for safety-critical LLM deployments beyond 6G, such as autonomous vehicles or medical triage.","The results are zero-shot; a testable extension is whether the 1B–1.5B transition persists under few-shot prompting or chain-of-thought, where smaller models sometimes close the gap to larger ones."],"forward_implications":["If the 1–1.5B stability transition is real, edge deployments should target 1.5–3B models for stability-critical intent translation, agentic coordination, and cross-slice arbitration rather than the largest model that fits.","Ultra-compact models (≤350M) remain useful only for low-risk, localized decisions; their large instability gap disqualifies them from safety-critical control loops.","The diminishing returns beyond 3B (+0.064 from 3B to 7B) imply that 7B-class models should be reserved for centralized orchestration where marginal robustness justifies higher latency and memory cost.","Domain-level differences mean trust and security tasks saturate earlier and can be served by smaller models, while intent and policy reasoning and distributed-intelligence tasks need more capacity.","The Edge Score ranking suggests a tiered deployment hierarchy, placing different model sizes at different points between radio/edge nodes and centralized control."],"fun_headline_variants":["Mid-size LLMs beat giants for 6G reasoning","1.5B models: the 6G reasoning sweet spot","Reasoning in 6G: small models, big efficiency","Edge AI: 1.5B models win on reliability per watt","Scaling law shift: mid-size LLMs dominate 6G edge"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central claims rest on two measurements: the instability gap Δ5 is computed from pass@5 on only 7 of 30 tasks while pass@1 covers all 3,722 questions, and the latency/VRAM profiles come from single-GPU bf16 inference that may not reflect real edge costs; if either is off, the 1–1.5B stability threshold and the Edge Score ranking shift.","fun_headline_variants_meta":{"raw":{"variants":["Mid-size LLMs beat giants for 6G reasoning","1.5B models: the 6G reasoning sweet spot","Reasoning in 6G: small models, big efficiency","Edge AI: 1.5B models win on reliability per watt","Scaling law shift: mid-size LLMs dominate 6G edge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1261,"prompt_tokens":987,"completion_tokens":274,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":731,"completion_tokens_details":{"reasoning_tokens":184}},"tokens_in":731,"tokens_out":274,"duration_ms":2779,"temperature":1.0,"reasoning_tokens":184,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T19:25:20.093248+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute A1 and A5 on the same 7-task subset used for pass@k and on the full 30-task set; if the 1B-to-1.5B Δ5 collapse disappears on the full set, the stability transition is a subset artifact. Separately, rerun the Edge Score with memory footprint measured at matched sequence lengths and with quantization enabled, and check whether mid-scale models still dominate the frontier.","supporting_citations":[],"review_version":1}