{"id":"28b81301-6c9c-46d1-865a-e318b84b4092","arxiv_id":"2605.27673","paper_version":1,"verdict":"ACCEPT","confidence":"LOW","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Complex-valued networks show task-dependent gains over real baselines on phase-sensitive data like PSK but not QAM, with large benchmark gaps often caused by hyperparameter instability rather than inherent superiority.","lead":"This paper compares complex-valued neural networks to several real-valued baselines on radio signal, quantum, and EEG tasks, finding that complex models help mainly on phase-sensitive problems but the advantage shrinks dramatically with proper tuning. A smart generalist might read it to avoid over-investing in complex architectures when real models suffice under matched conditions.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Whether 16-trial per-family search plus LR×activation factorial suffices to attribute RadioML gap to hyperparameter effects","rationale":"The reader's weakest_assumption directly identifies the load-bearing empirical claim; no more fundamental internal inconsistency (e.g., in the representation experiments or gradient analysis) is apparent from the provided text.","tokens_in":1886,"tokens_out":267,"duration_ms":13698,"concrete_test":"Re-run the real baseline families with a 100-trial random or Bayesian search over the same hyperparameter ranges plus AdamW, cosine scheduling, and two additional initializations; if best real accuracy rises by >3 PP relative to the reported 2.46 PP gap, the artifact explanation weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that the 22.94 PP gap is primarily an optimization artifact rests on the 16-trial independent tuning (plus the factorial) having adequately sampled the real-valued loss landscape. If real models possess better optima reachable only with wider ranges, different optimizers, or initializations outside the explored space, the attribution to 'high-learning-rate first-step instability' and 'complex parameter coupling' would not hold; the residual 2.46 PP gap could still reflect representational differences rather than tuning failure.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript claims that complex-valued neural networks (CVNNs) provide advantages in specific tasks involving phase or magnitude information but are not universally superior to real-valued models. This is demonstrated through representation-first evaluations on synthetic RF tasks (PSK, QAM, mixed), quantum wavefunction prediction, and EEG analytic-signal experiments. On the RadioML 2018.01A benchmark, an apparent 22.94 percentage point advantage for a CReLU complex model under matched-shared-trial selection reduces to 2.46 PP under independent per-family tuning with a 16-trial search space. Gradient analysis and a learning-rate × activation factorial experiment attribute the initial gap to hyperparameter-driven optimization instability in real baselines rather than representational superiority.","tokens_in":1977,"tokens_out":497,"duration_ms":27882,"significance":"If the results hold, the paper provides a valuable nuanced perspective on CVNN utility, emphasizing inductive biases tied to representation and symmetry rather than blanket superiority. Strengths include the use of multiple matched baselines (Cartesian real, polar, phase-only, magnitude-only, parameter-matched, FLOP-matched), FLOP-matched controls, gradient tracing, and the factorial hyperparameter experiment, which directly support claims about task dependence and benchmarking artifacts. This could guide future work in domains like RF, quantum, and neuroscience signals.","major_comments":[{"comment":"The attribution of the RadioML gap primarily to hyperparameter effects (reducing from 22.94 PP to 2.46 PP under independent per-family tuning) rests on the 16-trial search space plus LR×activation factorial sufficiently exploring the real-valued loss landscape. If better optima for real baselines exist outside this space (e.g., wider ranges or different optimizers), the residual gap may reflect representational differences rather than tuning failure. This is load-bearing for the benchmarking-artifact conclusion in the RadioML section.","section":"RadioML benchmarking experiment"}],"minor_comments":[{"comment":"The abstract clearly summarizes the claims but could briefly note the total number of non-RF domains evaluated to emphasize breadth.","section":"Abstract"},{"comment":"Notation for complex activations (e.g., CReLU) should be defined once in a dedicated subsection and referenced consistently in all experimental descriptions.","section":"Methods/Notation"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and the recommendation of minor revision. The single major comment concerns the sufficiency of the hyperparameter search in the RadioML experiment; we respond directly below.","responses":[{"response":"We agree that a finite search cannot guarantee that the global optimum for real-valued models has been found, and therefore cannot exclude the possibility that a residual gap after independent tuning reflects representational differences. The 16-trial per-family search and the LR×activation factorial were chosen specifically to target the high-learning-rate first-step instability identified in the gradient analysis; this design isolates a concrete optimization pathology rather than attempting exhaustive coverage. The observed collapse of the gap under these controls supports hyperparameter sensitivity as the dominant factor in the original matched-shared-trial comparison, yet we accept that the experiment does not constitute proof against all possible real-valued configurations. We will revise the RadioML section to state this limitation explicitly and to qualify the benchmarking-artifact claim accordingly.","revision_made":"partial","referee_comment":"[RadioML benchmarking experiment] The attribution of the RadioML gap primarily to hyperparameter effects (reducing from 22.94 PP to 2.46 PP under independent per-family tuning) rests on the 16-trial search space plus LR×activation factorial sufficiently exploring the real-valued loss landscape. If better optima for real baselines exist outside this space (e.g., wider ranges or different optimizers), the residual gap may reflect representational differences rather than tuning failure. This is load-bearing for the benchmarking-artifact conclusion in the RadioML section."}],"tokens_in":1525,"tokens_out":377,"duration_ms":20102,"standing_objections":["Exhaustively enumerating all hyperparameter ranges, optimizers, and architectures to prove that no superior real baseline exists is computationally intractable and lies outside the scope of the present study."]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that complex networks are not broadly superior; they help when the signal lives in phase, as in PSK tasks, while QAM favors magnitude-based real models and mixed cases show only small edges. Rotations break coordinate-dependent approaches without augmentation. The same split appears in quantum wavefunctions, where phase recovers momentum invisible to magnitude, and in EEG where phase locking, bursts, and coupling each favor different views.\n\nThe work does a clean job with its representation-first design. Matched real, polar, phase-only, magnitude-only, parameter-matched, and FLOP-matched baselines plus gradient tracing give direct support for the task-dependence claim. The RadioML section is the clearest addition: the 22.94 PP gap under shared trials shrinks to 2.46 PP with independent per-family tuning, and the LR-by-activation factorial plus first-step instability analysis traces it to real-model optimization fragility rather than representation.\n\nThe stress-test worry about the 16-trial space is reasonable but does not land as a central flaw here. The controls shown make the hyperparameter attribution hold for the tested regimes; a wider search might narrow the residual gap further but would not erase the pattern they document.\n\nThe main limitation is the purely empirical scope with no derivation of when the inductive bias should pay off. Domains stay narrow too.\n\nThis is useful for people building models on RF, quantum, or biosignal data who need to decide on complex arithmetic. It deserves a serious referee because the controls are tight and the artifact identification is new enough to matter for future benchmarking.","headline":"The paper shows CVNN gains are task-dependent on phase vs magnitude and that the big RadioML gap was mostly a hyperparameter artifact from real baseline tuning.","tokens_in":2432,"tokens_out":396,"would_cite":true,"duration_ms":25017,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Complex-valued networks help only when signals encode information in phase or magnitude-phase coupling, with RadioML gaps mostly from unequal tuning.","keywords":["complex-valued neural networks","representation geometry","RF signal classification","phase information","optimization stability","benchmarking artifacts","inductive bias"],"falsifier":"Finding a broader hyperparameter search or different optimizer that raises real baseline accuracy on RadioML 2018.01A to within roughly 3 percentage points of the complex model would show the reported gap is not primarily hyperparameter-driven.","tokens_in":2779,"feed_emoji":"📡","tokens_out":841,"duration_ms":28258,"temperature":0.7,"pith_summary":"The paper tests complex-valued neural networks against multiple real-valued baselines that vary in coordinate system and parameter count on RF, quantum, and EEG tasks. It establishes that complex models are not universally superior: phase-shift keying tasks reward phase-aware or complex representations while quadrature amplitude modulation tasks are better served by magnitude-only real models, and mixed signals yield only modest complex advantages. A central result is that the large RadioML 2018.01A gap between CReLU complex models and real baselines shrinks from 22.94 to 2.46 percentage points once each family receives its own hyperparameter search, because complex parameter coupling stabilizes gradients against high initial learning rates that destabilize real models. This matters because it shows CVNNs function as geometry-specific inductive biases whose value must be checked against task structure and fair optimization rather than assumed by default.","feed_headline":"Complex nets' RadioML lead drops from 23 to 2 points with matched tuning","feed_subtitle":"Gains appear only when tasks use phase or phase-magnitude coupling; otherwise real magnitude models suffice and tuning explains most reporte","key_machinery":"Side-by-side evaluation of Cartesian real, polar, phase-only, magnitude-only, parameter-matched real, and FLOP-matched real baselines, together with gradient analysis of loss-signal distribution through complex parameter coupling.","core_discovery":"Complex-valued neural networks are structured inductive biases whose effectiveness depends on alignment between data geometry and representation choice. On synthetic RF tasks, PSK-only signals favor phase-aware and complex-valued models, QAM-only signals favor magnitude-based models, mixed PSK+QAM yields only a small complex advantage, and unseen carrier-phase rotations degrade coordinate-dependent models without augmentation. Parallel patterns appear in quantum wavefunction prediction, where phase recovers momentum invisible to magnitude alone, and in EEG analytic signals, where phase locking, amplitude bursts, and phase-amplitude coupling each favor different coordinate views. On RadioML 2","pith_inferences":["Practitioners should run equivalent tuning budgets across real and complex families before crediting gains to complex arithmetic.","The conditional benefit pattern is likely to appear in other phase-sensitive domains such as audio or radar, where magnitude versus phase encoding can be isolated.","Whether complex coupling confers similar first-step stability under optimizers other than those tested remains open and directly testable."],"forward_implications":["PSK-only tasks favor phase-aware and complex-valued models over magnitude-only real models.","QAM-only tasks favor magnitude-based real models over phase-aware or complex ones.","Mixed PSK+QAM tasks produce only a small complex-valued advantage.","Unseen carrier-phase rotations break performance of coordinate-dependent models unless the training data includes augmentation.","The RadioML performance gap is primarily an artifact of unequal hyperparameter sensitivity rather than an inherent representational superiority of complex arithmetic."],"fun_headline_variants":["Matched tuning cuts complex RadioML lead to 2 points","Phase signals favor complex models over magnitude real baselines","CVNN gains depend on data geometry and coordinate choice","Quantum momentum needs phase not magnitude alone"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 16-trial per-family search space plus the learning-rate times activation factorial is assumed to have sufficiently explored the real baseline optimization landscape.","fun_headline_variants_meta":{"raw":{"variants":["Matched tuning cuts complex RadioML lead to 2 points","Phase signals favor complex models over magnitude real baselines","CVNN gains depend on data geometry and coordinate choice","Quantum momentum needs phase not magnitude alone"]},"model":"grok-4.3","cost_usd":0.006801,"raw_usage":{"total_tokens":3242,"prompt_tokens":828,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":68012000,"prompt_tokens_details":{"text_tokens":828,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2356,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":828,"tokens_out":58,"duration_ms":19418,"temperature":1.0,"reasoning_tokens":2356,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T18:43:01.433418+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding a broader hyperparameter search or different optimizer that raises real baseline accuracy on RadioML 2018.01A to within roughly 3 percentage points of the complex model would show the reported gap is not primarily hyperparameter-driven.","supporting_citations":[],"review_version":1}