{"id":"1646be0a-dc67-45c7-a1ea-076244b9e638","arxiv_id":"2607.06456","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A hardware-aware open-source SNN simulator embeds FG and ReRAM nonlinearities and multiple analog neuron models into training and reports accuracy plus area, power, and quantization metrics on neuromorphic benchmarks.","lead":"An open-source PyTorch framework trains mixed-signal spiking networks on physical floating-gate and ReRAM parameters instead of abstract weights, while reporting accuracy, area, and power. It helps designers pick neuron–synapse pairs that fit edge energy and silicon budgets before tape-out.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The predictive claim for design-space exploration rests on mean-behavior models whose fidelity is only validated at single-neuron ISI and tiny XOR scale, not at full-network accuracy/area/power rankings.","rationale":"The Reader correctly isolates the fidelity gap between mean-behavior Python models and deployable mixed-signal hardware as the weakest assumption. My stress-test simply sharpens the same point: the only circuit-level cross-checks are single-neuron ISI curves and a 2-8-2 XOR toy network; every accuracy/area/power number used for design-space claims (Tables 2–4) is generated inside the unvalidated full-network Python abstraction. Because the authors already list noise, mismatch, parasitics, AER latency and calibration as future work (§8), the concern is internal rather than external. No stronger objection (e.g., mathematical inconsistency or fabricated data) appears. The concrete Cadence re-implementation of one Table-2 network is the minimal experiment that would either confirm or overturn the predictive ranking claim. Hence the Reader’s CONDITIONAL verdict and high-confidence assessment remain appropriate; no change of category is warranted.","tokens_in":22189,"tokens_out":706,"duration_ms":7928,"concrete_test":"Take the smallest full-network configuration that still appears in Table 2 (e.g., SHD 700\to150\to20 FG+AH or N-MNIST 578\to250\to100\to10 ReRAM+LIF). Re-implement it at transistor level (or with Monte-Carlo mismatch + measured noise models) in Cadence/Verilog-A for a statistically meaningful subset of synapses/neurons; recompute test accuracy, total area and average power. If any ranking among the three synapse/neuron pairs reverses or any accuracy shifts by >5–7 % relative to the Python numbers, the design-space-exploration claim is not yet hardware-predictive.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that embedding calibrated FG/ReRAM nonlinearities and mixed-signal neuron models into PyTorch yields accuracy + area + power numbers that support quantitative cross-layer design-space exploration for mixed-signal SNNs. That claim is load-bearing on the assumption that mean-behavior, Euler-discretized, pulse-driven point-neuron models (fitted to I–V and ISI curves) produce rankings that survive real silicon. The paper itself supplies only two fidelity anchors: (1) single-neuron ISI vs. Cadence (Fig. 9, Fig. 10, Supp. S3) and (2) a 2\times8\times2 XOR network whose Cadence vs. Python area/power match to ~0.1 % and whose accuracy is 100 % (Sec. 6.1). Full-network results (Tables 2–4) are pure Python; no multi-neuron Cadence co-simulation, no injected mismatch/noise, and no AER/routing is present. Section 8 and the supplementary temporal-deviation analysis explicitly flag these omissions. Consequently the reported 8–11 % accuracy drops and the 100\times area advantage of ReRAM vs. FG are simulation-internal trade-offs whose ordering may reverse once device variation, parasitics, or event-routing latency are included. The leap from mean-model numbers to “hardware-predictive” exploration is therefore the softest link.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript presents an open-source PyTorch framework for mixed-signal SNN design-space exploration that embeds experimentally calibrated floating-gate (65 nm) and ReRAM (Skywater 130 nm) nonlinear synapse models, together with multiple mixed-signal neuron models (Axon-Hillock, adaptive/standard/Schmitt LIF, Hodgkin–Huxley), into end-to-end training. Synaptic parameters optimized during learning are physical device quantities (V_FG0, filament gap g) rather than abstract weights. The tool supports fully connected and recurrent architectures, is evaluated on N-MNIST, DVS Gesture and SHD, and reports classification accuracy jointly with estimated silicon area, power and post-training quantization sensitivity. A small 2×8×2 XOR network is cross-checked against Cadence (function, area, power), and single-neuron ISI curves are matched to transistor-level simulations.","tokens_in":22640,"tokens_out":1484,"duration_ms":19552,"significance":"If the framework’s accuracy–area–power rankings are accepted as indicative of mixed-signal hardware trade-offs, this is a practically useful contribution: it lowers the barrier between algorithmic SNN work and analog neuromorphic circuit constraints, ships open code, and unifies several calibrated device and neuron models under one training loop. Strengths include measurement-anchored FG and ReRAM I–V models, Cadence-validated neuron ISI behavior, an explicit XOR circuit co-simulation, and joint reporting of accuracy with hardware metrics on standard neuromorphic datasets. The main value is as a co-design exploration platform rather than as a new learning algorithm or a fully predictive silicon model.","major_comments":[{"comment":"The central claim that the framework supports quantitative, hardware-predictive design-space exploration (Abstract; §1; §7) rests on mean-behavior models whose network-level fidelity is only weakly validated. Single-neuron ISI matches Cadence (Fig. 9–10; Supp. S3) and a 2×8×2 XOR net matches area/power to ~0.1% with 100% function (§6.1). Full-network results in Tables 2–4 are pure Python: no multi-neuron Cadence co-simulation, no injected mismatch/noise/parasitics/drift, and no AER/routing. §8 correctly flags these gaps, but the abstract and conclusion still present accuracy drops and ReRAM vs FG area/power rankings as hardware-oriented design guidance. Either temper the claim to “simulation-internal, mean-model exploration” or add at least one multi-neuron fidelity check (e.g., mismatch Monte Carlo or a larger Cadence/Python co-sim) that shows rankings are stable.","section":null},{"comment":"Table 3 and §6: area and power are obtained by analytically scaling Cadence core-block layouts and accumulating branch/event currents × V_DD (Eq. 16). This omits interconnect, AER arbitration, memory access, and array parasitics that often dominate mixed-signal neuromorphic chips. The reported >100× area advantage of ReRAM vs FG and the absolute mW-scale numbers are therefore not yet comparable to deployable systems. Please state explicitly what is included/excluded in the area/power model, and either (i) add a sensitivity analysis under plausible routing/peripheral overheads or (ii) reframe Table 3 as core-compute estimates only, not full-chip metrics.","section":null},{"comment":"§6.2 / Discussion: accuracy degradations of ~8–11% (N-MNIST) and ~7–8% (DVS) when replacing nn.Linear with ReRAM/FG are attributed to “hardware-induced non-idealities,” yet the authors note that disentangling nonlinearity, limited precision, and architecture is left for future work. Without an ablation (ideal linear weights vs nonlinear continuous device model vs post-training quantization alone; Tables 2 vs 4), the design-space conclusions about which neuron–synapse pair “best satisfies” accuracy–energy–area constraints are under-supported. A minimal ablation on one dataset would substantially strengthen the comparative claims.","section":null},{"comment":"§3.1 and Table 4: quantization is applied post-training (8-bit FG, 3-bit ReRAM) by nearest-level projection after continuous optimization of V_FG0 or g. The framework’s stated advantage is optimizing physical parameters inside the training loop; discrete programmable levels are not part of that loop. For ReRAM especially (only 8 levels), PTQ can dominate the accuracy gap. Either integrate differentiable or STE-based quantization during training, or clearly separate “device-nonlinear continuous training” from “hardware-level discrete programming” when interpreting Table 4.","section":null}],"minor_comments":[{"comment":"Table 1 comparison is thin on modern hardware-aware or mixed-signal SNN tools (e.g., Brian2/Lava with device plugins, NeuroSim-class CIM estimators, SANA-FE cited only in limitations). A short positioning paragraph would help readers place the contribution.","section":null},{"comment":"Fig. 1 caption and body: “eThe implementation” appears to be a typo; also “SnnTorch” / “snnTorch” capitalization is inconsistent.","section":null},{"comment":"Eq. (2) text refers to Q_FG but the displayed equation uses capacitive coupling terms only; align notation with the prose.","section":null},{"comment":"Table 2: architecture strings and neuron process nodes (65 nm vs 28 nm) are useful; please also report number of trainable device parameters and whether recurrent weights use the same FG/ReRAM model as feedforward weights.","section":null},{"comment":"§5.1.2: Mutual SNN Pooling is important for DVS results; a one-sentence statement of whether pooling is fixed or learned, and whether it is counted in area/power, would avoid ambiguity.","section":null},{"comment":"Code link (§9) is welcome; please pin a commit/tag and list which tables/figures are fully reproducible from the public repo.","section":null},{"comment":"Supplementary temporal-deviation analysis (Fig. S3) is valuable; consider promoting a short quantitative summary (median |ΔV|, ISI error) into the main text near Fig. 9–10.","section":null}],"recommendation":"major_revision","confidential_remarks":"The work is a solid engineering contribution and fits eess.SP / neuromorphic design audiences, but the title and abstract currently oversell “hardware-aware” predictive power relative to the validation provided. I would accept after the authors either narrow the claims or add a modest network-level fidelity/ablation study; I would not reject on novelty grounds. Open-source intent and measurement-calibrated models are genuine positives for the community."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a practical engineering paper that ships something the mixed-signal neuromorphic community actually needs: an open PyTorch framework that trains directly on physical FG (V_FG0) and ReRAM (filament gap) parameters, not abstract weights, with several Cadence-aligned neuron models (AH, adaptive/standard/Schmitt LIF, HH) across 28/65 nm and joint accuracy–area–power–quantization numbers on N-MNIST, DVS Gesture, and SHD.\n\nWhat is new is the unification, not the individual pieces. Device equations are fitted to measured FG 65 nm and Skywater 130 nm ReRAM I–V; neuron ISI curves match Cadence; a 2\times8\times2 XOR is cross-checked in Cadence (100 % function, area/power agreement to ~0.1 %). Tables 2–4 then show the expected accuracy drop when ideal synapses are replaced by nonlinear devices, plus the large area/power advantage of ReRAM over FG and the effect of 3-bit vs 8-bit quantization. Code is promised. That is real, usable work.\n\nThe soft spot is exactly where the stress-test points, and the authors already flag it in §8. Full-network accuracy/area/power rankings live entirely in mean-behavior Python models. There is no multi-neuron Cadence co-simulation, no mismatch/noise/parasitics/drift, and no AER routing. Single-neuron ISI and one tiny XOR do not guarantee that the reported 8–11 % accuracy gaps or the 100\times area ranking will survive real silicon. That does not make the tool worthless; it makes the “hardware-predictive design-space exploration” claim overstated relative to the fidelity evidence. Circularity is low—losses are external classification losses and hardware metrics come from separate layout baselines.\n\nMath and citations look solid for an methods paper; self-cites to their own FG/ReRAM characterization are the measurement anchors, not circular. Free parameters (device fits, neuron currents, bit-widths, architecture) are normal for this genre.\n\nWho it is for: people building or evaluating mixed-signal SNN hardware who want a starting co-design sandbox. It deserves a serious referee. I would engage with the code and cite the framework when I need a baseline that already folds FG/ReRAM nonlinearities into training; I would not treat the full-network rankings as silicon truth without the missing non-ideality studies.","headline":"Useful open mixed-signal SNN co-design tool with real device-parameter training; the hardware-predictive claim is only weakly anchored beyond single-neuron ISI and a tiny XOR net.","tokens_in":23275,"tokens_out":600,"would_cite":true,"duration_ms":6210,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"An open-source framework embeds real floating-gate and ReRAM device physics into SNN training so designers can optimize physical synaptic parameters and jointly score accuracy, area and power.","keywords":["mixed-signal SNNs","hardware-aware training","floating-gate synapses","ReRAM","neuromorphic design space exploration","spiking neural networks","edge computing","quantization sensitivity"],"falsifier":"Build a small mixed-signal SNN with the same FG or ReRAM synapses and neuron circuits, load the tool-optimized physical parameters with no extra tuning, and check whether measured classification accuracy, power and area on one benchmark match the framework's reported numbers within the claimed margins.","tokens_in":23086,"feed_emoji":"⚡","tokens_out":868,"duration_ms":17498,"temperature":0.7,"pith_summary":"Energy-efficient edge intelligence needs simulators that capture how analog synapses and neurons actually behave, not only ideal spike models. This paper presents an open-source hardware-aware framework that places experimentally calibrated floating-gate and ReRAM synapse equations, plus several mixed-signal neuron circuits, inside custom PyTorch modules. Training therefore updates physical parameters such as floating-gate voltage or filament gap rather than abstract digital weights. On N-MNIST, DVS Gesture and Spiking Heidelberg Digits the tool returns classification accuracy together with estimated silicon area, power and quantization sensitivity, letting designers compare neuron, synapse and architecture choices against concrete hardware constraints before fabrication.","feed_headline":"Train SNNs on real chip physics, not abstract weights","feed_subtitle":"Open tool optimizes floating-gate and ReRAM parameters while scoring accuracy, area and power.","key_machinery":"Custom PyTorch modules that implement the modified EKV floating-gate current law and the exponential-sinh ReRAM I-V characteristic so the trainable weights are the device parameters V_FG0 and filament gap g; surrogate-gradient BPTT then updates those physical parameters while Euler-discretized neuron models (Axon-Hillock, LIF variants, Hodgkin-Huxley) produce spikes.","core_discovery":"By embedding nonlinear, measurement-calibrated floating-gate and ReRAM synapse models and multiple mixed-signal neuron implementations directly into a PyTorch training and inference loop, the framework enables end-to-end optimization of physical synaptic parameters and quantitative cross-layer design-space exploration of accuracy, area, power and quantization effects on standard neuromorphic benchmarks.","pith_inferences":["If mean-device models prove sufficient, co-training on physical parameters could cut the lengthy bias and weight calibration that currently limits mixed-signal neuromorphic chips.","Adding device-noise and mismatch distributions to the same embedding would turn the explorer into a variability-aware optimizer rather than a mean-behavior tool.","An open common harness of this form could become a shared benchmark for comparing future analog synapse technologies on identical SNN tasks and metrics."],"forward_implications":["Designers can rank floating-gate versus ReRAM and LIF versus Axon-Hillock or Hodgkin-Huxley configurations by joint accuracy-area-power metrics before tape-out.","Post-training 8-bit floating-gate and 3-bit ReRAM quantization can be scored for accuracy drop without a separate abstract-to-device mapping step.","Recurrent and fully connected architectures can be compared under identical hardware models on temporal tasks such as SHD.","Neuron parameters extracted from both 65 nm and 28 nm processes can be swapped inside the same training pipeline, supporting multi-node exploration."],"fun_headline_variants":["Train mixed-signal SNNs on real device physics in PyTorch","Hardware-aware open tool explores SNN neurons and analog synapses","Optimize floating-gate and ReRAM params, not abstract weights","Cross-layer SNN design: accuracy vs area power and fidelity","Mixed-signal SNN framework reports silicon metrics on neuromorphic tasks"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That average-behavior device models and simple point-neuron circuits fitted to measured I-V and spike-rate curves are faithful enough to predict real silicon accuracy, area and power without modeling noise, mismatch, parasitics, routing delays or on-chip calibration.","fun_headline_variants_meta":{"raw":{"variants":["Train mixed-signal SNNs on real device physics in PyTorch","Hardware-aware open tool explores SNN neurons and analog synapses","Optimize floating-gate and ReRAM params, not abstract weights","Cross-layer SNN design: accuracy vs area power and fidelity","Mixed-signal SNN framework reports silicon metrics on neuromorphic tasks"]},"model":"grok-4.5","effort":"low","cost_usd":0.001966,"raw_usage":{"total_tokens":839,"prompt_tokens":767,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":19660000,"prompt_tokens_details":{"text_tokens":767,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":0,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":767,"tokens_out":72,"duration_ms":1438,"temperature":1.0,"reasoning_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T16:00:56.092491+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Build a small mixed-signal SNN with the same FG or ReRAM synapses and neuron circuits, load the tool-optimized physical parameters with no extra tuning, and check whether measured classification accuracy, power and area on one benchmark match the framework's reported numbers within the claimed margins.","supporting_citations":[],"review_version":2}