{"id":"97097c2d-874b-4578-abc4-de65fddeaa6e","arxiv_id":"2506.18242","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Two trainable metasurface layers, used in opposite orders, classify MNIST digits and FashionMNIST items with one shared set of hardware.","lead":"This paper builds an optical neural network from two stacked metasurface lenses that can switch between recognizing handwritten digits and recognizing clothing by simply swapping the order of the two layers. The authors show in simulations and a terahertz experiment that the same hardware handles both tasks with accuracy close to separate networks while using half the optical elements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is under-tested: no ablation isolates whether layer-order permutation, rather than the shared phase pair or weight sharing, actually causes the multi-task capability.","rationale":"The reader's weakest assumption focuses on simulation-to-experiment fidelity of the metasurface forward model, which is a real concern but is a known and partially acknowledged limitation (Sec. 4 mentions forward-design errors and future inverse design). The experiment does demonstrate both tasks physically, and the claim is explicitly about the architecture and training method, not about device-grade accuracy. The more load-bearing, correctness-relevant concern is whether the simulation itself isolates the permutation mechanism as the cause of multi-task capability. The paper provides no ablation that separates (a) order permutation, (b) shared-layer multitask learning with fixed order, and (c) the plain expressive power of the two-layer phase pair. Without such controls, the central claim that layer rearrangement is the reconfiguration mechanism is under-determined by the evidence. I therefore disagree with the reader's choice of weakest assumption, while agreeing with the overall CONDITIONAL verdict because the concern is testable and not fatal.","tokens_in":10632,"tokens_out":1643,"duration_ms":17090,"concrete_test":"Retrain the two-task A-DNN under three controlled conditions with identical hyperparameters and at least five random seeds: (A) the actual weighted loss of Eq. 7 with order {L1,L2} for MNIST and {L2,L1} for FashionMNIST; (B) a control where both tasks are trained with the same order {L1,L2} under the same weighted loss, testing whether the shared phase pair can still classify both tasks (this quantifies implicit multiplexing without permutation); and (C) a control where the two layers are constrained to be identical (single phase profile duplicated), testing whether layer-order diversity is essential. If (B) or (C) achieves accuracy comparable to (A), the permutation mechanism is not load-bearing; if (A) clearly outperforms both, the central claim is supported. Also report per-seed mean and standard deviation to rule out initialization artifacts.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that the same two phase-only metasurface layers, when swapped, yield two distinct trained classifiers (Sec. 3.2, Table 1). This requires that the optimized solutions for the two tasks are not accidentally identical under reversal, and that the task separation genuinely comes from layer order rather than from implicit multiplexing or from one task being solved by the second layer with the first acting as a weak perturbation. The paper reports only one training run per condition, with no repetition across seeds or hyperparameters, so the 88.9%/81.8% results could reflect a favorable initialization. More importantly, no ablation (e.g., fixing the two layers to identical phase profiles, training a single pair without the swap-sharing constraint, or comparing with a control where both tasks use the same order) is reported to confirm that order permutation, rather than the pair together, carries the multi-task capability. The robustness analysis (Sec. 3.3) shows tolerance to errors but does not test the sensitivity of the swap mechanism to layer similarity; if the two trained phase profiles are highly correlated, the mechanism may be fragile. The experimental demonstration (Sec. 3.4) uses a forward-design model and reports a simulation-to-experiment drop (93.1% to 75% for digits, 87.2% to 70% for fashions), which the reader already flagged; but even in simulation the central claim is not fully isolated. Therefore, the most load-bearing concern is not fabrication fidelity but the lack of evidence that the task-specific behavior is attributable specifically to layer-order permutation rather than to the joint phase pair under the weighted loss constraint.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an arrangeable diffractive neural network (A-DNN) in which two phase-only metasurface layers are trained jointly under a weighted multi-task loss (Eq. 7); at inference, swapping the physical order of the two layers switches the network between MNIST digit and FashionMNIST fashion classification. Simulations on 10-class tasks report test accuracies of 88.9% (MNIST) and 81.8% (FashionMNIST), versus 91.0% and 84.4% for two dedicated two-layer D2NNs, using one shared pair of layers (Table 1). A 5-class proof-of-concept experiment at 0.291 THz with 80x80 metasurfaces reports simulated accuracies of 93.1% and 87.2% but measured accuracies of 75% and 70% on 40 fabricated input samples. The paper also presents a beta-sweep showing task-weight control and robustness tests against phase noise and layer misalignment.","tokens_in":10924,"tokens_out":8711,"duration_ms":102328,"significance":"The central idea, if confirmed, is attractive: layer-order permutation is a zero-hardware-cost reconfiguration mechanism that turns one trained diffractive pair into two classifiers, and the weighted-loss formulation gives a simple way to trade off task performance. The manuscript has real strengths: results are measured on standard external benchmarks; the beta trade-off is swept openly; the training code is promised to be available; and a full metasurface fabrication and THz measurement is reported. The significance is nevertheless moderate: the simulations show a small but real accuracy penalty relative to dedicated D2NNs, the experimental demonstration rests on 40 samples with a large simulation-to-experiment drop, and the mechanism is not isolated by ablations.","major_comments":[{"comment":"The central claim that swapping layer order creates two distinct classifiers is not isolated by the experiments. Training under Eq. (7) optimizes a shared phase pair against both orderings simultaneously, so the same pair could in principle solve both tasks without the swap being causally important (e.g., one layer could dominate each task, or the pair could multiplex the two tasks in a way that is insensitive to order). I request (i) a control in which the same two layers are trained on both tasks with a fixed order (using a combined output plane) and evaluated without any reconfiguration, (ii) a control in which a single-task-trained pair is evaluated after swapping its layers on the other task, (iii) a similarity measure between the optimized phase profiles, and (iv) training from at least five random initializations with mean and standard deviation for every reported accuracy. Without (i)-(iv), the 88.9%/81.8% results could reflect a favorable initialization or an implicit multiplexing effect rather than the permutation mechanism.","section":"Sec. 3.2, Eq. (7)"},{"comment":"The experimental demonstration does not support the quantitative accuracy comparison. The text says 40 samples were fabricated but does not state the per-task split. For a five-class task, chance is 20%, and 30 correct out of 40 gives a 95% Wilson interval of roughly 59-86%, so the reported 75% and 70% values are not precise estimates of deployment accuracy. The simulation-to-experiment drop (93.1% to 75% for digits, 87.2% to 70% for fashions) is attributed to forward-design errors (Sec. 4, Fig. S8) without a quantitative error budget. Please report per-class counts, confidence intervals, and either a larger experimental test set or a simulation that uses measured meta-atom phase and amplitude responses to reproduce the experimental accuracies.","section":"Sec. 3.4, Figs. 7-8"},{"comment":"The claims of '50% improvement in hardware efficiency' and 'saving nearly half of the training time' are not demonstrated by the reported measurements. Table 1 compares parameter counts, but A-DNN's per-iteration forward cost includes two ordered propagations (one per task), and no wall-clock training time or energy measurement is reported. The comparison also varies the number of tasks by construction (one multi-task network versus two single-task networks), so the claimed savings need a direct timing comparison with the same optimizer and hardware. Please either report actual training times or energy or restrict the claim to parameter count.","section":"Sec. 3.2, Table 1"},{"comment":"The robustness analysis does not test the sensitivity of the permutation mechanism itself. The Gaussian phase-noise and misalignment tests show that the trained A-DNN tolerates fabrication errors, but they do not address whether the classification of each task remains tied to the correct layer order, nor how the mechanism degrades as the two phase profiles become more correlated. A direct test would be to perturb one layer and measure the accuracy of both orderings, or to interpolate between the two optimized phase profiles and report the accuracy of each task; this would show whether the swap is a true functional switch rather than a fortuitous property of the optimized pair.","section":"Sec. 3.3"}],"minor_comments":[{"comment":"Eq. (1) presents a Rayleigh-Sommerfeld propagator, while Fig. 2 states that forward propagation is computed with the angular-spectrum method; clarify which model is implemented in the training code and confirm the two are used consistently.","section":"Sec. 2, Fig. 2"},{"comment":"The 'overall test accuracy' in Fig. 3d is not defined; specify whether it is the mean of the MNIST and FashionMNIST accuracies and add error bars across seeds before claiming that moderately increasing beta improves overall accuracy.","section":"Sec. 3.1, Fig. 3d"},{"comment":"The accuracies 93.1% and 87.2% are for five-class subsets, while Table 1 reports 88.9% and 81.8% for ten-class tasks; state the class count at every occurrence to avoid confusion.","section":"Sec. 3.4"},{"comment":"The code link is given as 'Atrf/Arrangeable-multi-task-diffractive-neural-network' without a full URL or DOI; provide a persistent link that reviewers can access.","section":"Data Availability"},{"comment":"The caption says x, y, and z represent the dimensions of the diffraction layer, while the text says x and y are side lengths and z is the layer distance; align the caption with the text.","section":"Fig. 5 caption and Sec. 3.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a sharper positioning against Refs. [39]-[44]; the novelty relative to P-DNN [41] and multiplexed single-layer metasurface approaches [43,44] is asserted but not quantified. If the authors cannot provide the ablations requested in the first major comment, the title claim should be narrowed accordingly. The experimental section is best described as a proof-of-concept; as written, the 40-sample validation is too weak to carry the quantitative claims in the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea is worth your time: train one pair of phase-only layers, then swap their order to get two different classifiers. That is a clean reconfiguration mechanism, and I don't see it in the cited prior work, which relies on external programmability or optical multiplexing. The simulation comparison against dedicated D2NNs shows a small accuracy gap (88.9/81.8 vs 91.0/84.4) while using half the diffractive units and one training run. The weighted multi-task loss is routine, but the paper uses it to set task priority, and the beta sweep in Fig. 3 is a reasonable demonstration. Credit also for shipping code and for openly attributing the sim-to-experiment drop to forward-design errors rather than hiding it.\n\nThe soft spots are real but not fatal to the concept. The most load-bearing issue is that no ablation isolates cause. You train one pair with a shared weighted loss and get two tasks; you don't show that layer order, rather than the joint phase pair, is what separates the tasks. A control with identical layers, or with a single task pair and no order swap, would pin the mechanism down. Also, only one training run per condition, so the 88.9/81.8 numbers could be a lucky initialization. The experimental demo is a proof-of-concept with 40 samples, no error bars, and a 75/70% accuracy that is well below simulation; that is fine for feasibility, but not enough to claim robust dynamic reconfiguration. The paper also says \"dynamic reordering\" when the demo is a manual swap, and there is no quantitative comparison with the closest prior reconfigurable or multiplexing D2NNs. Those are fixable with more analysis and clearer framing.\n\nThe stress-test note's central worry is valid: the mechanism is under-isolated. But I would not call it a load-bearing flaw that sinks the paper. The concept is credible, the simulations are consistent with it, and the authors are candid about the experimental gap. The right referee will ask for the missing ablations and more seeds, not desk-reject.\n\nWho is this for? People working on diffractive optical neural networks and THz metasurface hardware. It is a modest but real step toward reconfigurable optical classifiers. I would accept it for peer review and ask for a revision that isolates the order-permutation mechanism and adds basic statistical care to the experiment.","headline":"Layer-order permutation as a task selector is a genuinely new idea and the simulation evidence largely supports it, but the paper never isolates the mechanism and the experiment is too thin to carry the load.","tokens_in":11502,"tokens_out":1271,"would_cite":true,"duration_ms":17660,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By swapping the order of two phase-only metasurface layers, a single trained optical network classifies both handwritten digits and fashion items, at simulated accuracies close to two dedicated networks that use twice the hardware.","keywords":["optical neural networks","diffractive deep neural networks","multi-task learning","metasurfaces","terahertz","layer rearrangement","Pancharatnam-Berry phase"],"falsifier":"Measure the output-plane field of the two fabricated metasurfaces in both orders on the same input set: the permutation mechanism predicts that each ordering routes energy to a different class region with classification accuracy near the trained values, so statistically indistinguishable confusion matrices across orderings, or unchanged outputs when the layers are swapped with an unrelated trained pair, would refute the claim that layer order carries the reconfiguration.","tokens_in":10412,"feed_emoji":"🔄","tokens_out":12073,"duration_ms":113934,"temperature":0.7,"pith_summary":"This paper sets out to show that a diffractive neural network — an optical computer built from stacked phase-modulating layers — can be switched between tasks simply by reordering its already-designed layers, with no retraining and no new hardware. Using two phase-only metasurface layers trained once with a weighted multi-task loss, the authors demonstrate that the order $\\{\\varphi_1, \\varphi_2\\}$ classifies handwritten digits while the swapped order $\\{\\varphi_2, \\varphi_1\\}$ classifies fashion items, reaching simulated test accuracies of 88.9% and 81.8% against 91.0% and 84.4% for two dedicated two-layer networks. Because light encounters the same two modulators in a different sequence, each arrangement behaves as a distinct network, turning layer arrangement into a reconfiguration mechanism that costs nothing extra in fabrication. A proof-of-concept experiment at 0.291 THz confirms both orderings focus light onto the correct detection regions, though with lower accuracy than simulation. If the claim holds, static diffractive hardware can serve many tasks at a fraction of the hardware and training cost.","feed_headline":"Swapping layer order reconfigures one optical network for two tasks","feed_subtitle":"The same two phase-only metasurfaces, reordered per task, come close to dedicated networks at half the hardware cost.","key_machinery":"The load-bearing object is the ordered sequence of diffractive layers, treated as a permutation of shared phase-only metasurfaces. In a D2NN each layer multiplies the propagating optical field by its phase profile, and diffraction between layers is modeled with the angular spectrum method; since propagation from layer A to layer B is not identical to propagation from B to A, each ordering composes the same modulators into a different classifier. Training couples the orderings through a weighted multi-task loss, $\\mathcal{L}_{\\mathrm{multi}} = \\mathcal{L}(\\theta_1; X_1, Y_1) + \\beta\\mathcal{L}(\\theta_2; X_2, Y_2)$, so one backpropagation updates the shared phase profiles for both tasks and $\\beta$ trades accuracy between them. On the physical side, the layers are silicon metasurfaces whose rotation-angle-controlled Pancharatnam–Berry phase gives full $0$–$2\\pi$ phase coverage, designed by FDTD parameter sweeps and fabricated by photolithography, so one fabricated wafer pair realizes both task configurations by mechanically swapping layer order.","core_discovery":"The central claim is that permutation of trained diffractive layers is itself a task-selection mechanism. The authors construct a two-task A-DNN from two phase-only metasurfaces with phase profiles $\\varphi_1$ and $\\varphi_2$, define the two configurations $\\theta_1 = \\{\\varphi_1, \\varphi_2\\}$ and $\\theta_2 = \\{\\varphi_2, \\varphi_1\\}$, and train both at once by minimizing the weighted sum $\\mathcal{L}_{\\mathrm{multi}} = \\mathcal{L}(\\theta_1; X_1, Y_1) + \\beta \\mathcal{L}(\\theta_2; X_2, Y_2)$ in a single backpropagation pass. After training, the network scores 88.9% test accuracy on MNIST and 81.8% on FashionMNIST, versus 91.0% and 84.4% for two separately trained two-layer D2NNs — halving the trainable diffractive units and roughly halving training time. The weight $\\beta$ acts as an accuracy dial, moving performance toward FashionMNIST when increased and toward MNIST when decreased. A three-task variant built from four layers raises the hardware-efficiency gain to 66.7% and cuts training time to one-third of the dedicated-network baseline.","pith_inferences":["Only the two-layer, two-task case is demonstrated in the main text, so the promise that larger stacks host many permutations as usable tasks is an extrapolation; a natural test is training a three- or four-layer stack and evaluating all $N!$ orderings to see how many reach usable accuracy.","The simulation-to-experiment accuracy gap (93.1% to 75% for digits, 87.2% to 70% for fashions) indicates the forward optical model's fidelity, not the permutation idea itself, is the current bottleneck; the paper's own suggested remedies, repeating identical nano-fins per pixel and inverse design, could be tested directly by replacing ideal phase profiles with measured complex transmission in simu","If layer order is a genuine task selector, the number of tasks a stack can host grows factorially with layer count, but accuracy per ordering may degrade as the same modulators are asked to do more; measuring that accuracy-versus-permutations tradeoff would tell whether the mechanism scales beyond a few tasks."],"forward_implications":["Reconfiguration no longer requires refabrication or retraining: any permutation of a trained stack is a candidate new configuration, so a single fabricated set of metasurfaces can be physically reordered for different jobs.","Hardware cost shrinks with the task-to-layer ratio: two tasks on two shared layers use half the diffractive units of two dedicated networks, and four layers hosting three tasks save 66.7% of hardware while cutting training time to one-third.","The weight parameter $\\beta$ gives an operator an explicit accuracy dial, allowing a harder or more valuable task to be prioritized at the expense of another.","The authors note the framework composes with other optical multiplexing techniques, so layer arrangement could be combined with wavelength- or spatial-multiplexing to enlarge the number of tasks per device."],"supporting_citations":[{"why":"Supplies the baseline D2NN architecture and the physical diffraction model the A-DNN builds on.","marker":"[17]"},{"why":"The pluggable diffractive neural network (P-DNN), the prior reconfigurable-D2NN approach the A-DNN is positioned against.","marker":"[41]"},{"why":"Prior multi-task D2NN work using single-layer metasurface multiplexing, which A-DNN extends to cascaded layers.","marker":"[43]"},{"why":"Provides the two-layer cascaded phase-only metasurface platform used for the experimental implementation.","marker":"[45]"},{"why":"Source of the MNIST handwritten-digit dataset for the first task.","marker":"[46]"},{"why":"Source of the FashionMNIST dataset for the second task.","marker":"[47]"},{"why":"Standard angular-spectrum and Fourier-optics reference underlying the diffraction forward model used in training.","marker":"[48]"},{"why":"Introduces the geometric (Pancharatnam–Berry) phase used to encode the full 0 to 2π phase control in the metasurfaces.","marker":"[49]"}],"fun_headline_variants":["Layer reorder turns one metasurface pair into two-task network","Reorder metasurfaces to switch tasks without new hardware","One optical network, two tasks, just swap layer order","Weighted loss tunes task focus in reconfigurable optical neural net","Metasurface layers rearrange to cut hardware cost in half"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scheme assumes that the angular-spectrum, phase-only optical model used during training predicts how the fabricated metasurfaces actually transform light; the paper's own experiment shows accuracy dropping from 93.1% to 75% for digits and from 87.2% to 70% for fashions, so that fidelity is only partly met.","fun_headline_variants_meta":{"raw":{"variants":["Layer reorder turns one metasurface pair into two-task network","Reorder metasurfaces to switch tasks without new hardware","One optical network, two tasks, just swap layer order","Weighted loss tunes task focus in reconfigurable optical neural net","Metasurface layers rearrange to cut hardware cost in half"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1551,"prompt_tokens":1034,"completion_tokens":517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":434}},"tokens_in":650,"tokens_out":517,"duration_ms":5739,"temperature":1.0,"reasoning_tokens":434,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:23:20.607260+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the output-plane field of the two fabricated metasurfaces in both orders on the same input set: the permutation mechanism predicts that each ordering routes energy to a different class region with classification accuracy near the trained values, so statistically indistinguishable confusion matrices across orderings, or unchanged outputs when the layers are swapped with an unrelated trained pair, would refute the claim that layer order carries the reconfiguration.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the baseline D2NN architecture and the physical diffraction model the A-DNN builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The pluggable diffractive neural network (P-DNN), the prior reconfigurable-D2NN approach the A-DNN is positioned against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior multi-task D2NN work using single-layer metasurface multiplexing, which A-DNN extends to cascaded layers."},{"cited_title":"Lecun, L","cited_arxiv_id":null,"evidence_quote":"Source of the MNIST handwritten-digit dataset for the first task."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Standard angular-spectrum and Fourier-optics reference underlying the diffraction forward model used in training."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the geometric (Pancharatnam–Berry) phase used to encode the full 0 to 2π phase control in the metasurfaces."}],"review_version":1}