{"id":"d51fa48f-520f-4fb0-8f3b-4da09ee8863b","arxiv_id":"2602.04582","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Direct analog microphone signals injected into BrainScaleS-2 neuron circuits localize transient sounds via a Jeffress-model spiking network and drive a servo motor on-chip.","lead":"A neuromorphic chip takes analog audio directly from microphones—no analog-to-digital conversion—and uses a spiking neural network to find which direction a sound came from, then turns a servo motor to face it. It shows how accelerated analog neuromorphic hardware can close a sensor-to-actuator loop entirely on one chip.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"End-to-end evidence does not quantify real-room localization accuracy; the only quantified evaluation bypasses the microphones with artificial ITD shifts.","rationale":"The reader's weakest_assumption focuses on the equivalence between synthetic spikes and analog injection for Fig. 5B. That is a fair concern for the mechanistic figure, but it is partially mitigated by Fig. 5C/D, which use actual analog playback through the preprocessing board and show monotonic end-to-end performance. The more load-bearing gap is between the quantified playback condition (artificial ITDs, no acoustic propagation) and the claim of reliable localization in a real room. The manuscript itself flags the playback bypass, and the physical demo is not quantified. This supports the existing CONDITIONAL verdict: the core hardware integration is demonstrated, but the headline claim needs additional real-world evidence. I therefore keep the reader's verdict unchanged while identifying a different emphasis for the condition.","tokens_in":7696,"tokens_out":8104,"duration_ms":96180,"concrete_test":"Run the untouched on-chip pipeline with a loudspeaker or calibrated click source at known azimuths (e.g., -90° to +90° in 15° steps, distance ≥1 m) in the same room, recording the Algorithm 1 output for ≥100 trials per angle. Compute the mean absolute angular error and trial-to-trial spread, and compare with the sound-card playback ITD sweep from Fig. 5D. If the real-source error or spread substantially exceeds the 2-4 unit variation seen in Fig. 5D, or shows systematic bias, the real-room claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that the system can \"reliably predict the position of a sound ... in a room\" and align a servo to it. The only quantitative evidence for this (Fig. 5C/D) is explicitly obtained by bypassing the microphones: \"For an automated evaluation of system performance over multiple ITDs, we bypass the microphones and play back a recorded clap with an adjustable, software-set inter-channel delay\" (Section II-A). This exercises the analog input path and the network, but not the acoustic conditions of the claimed scenario: variable source azimuth/distance, room reflections, reverberation, background noise, and ear-dependent amplitude/spectral cues. The physical demonstrations with claps or a ping-pong ball are anecdotal, with no measured angles, trial counts, or error statistics. The synthetic-spike issue in Fig. 5B is real but less load-bearing, because Fig. 5C/D already show that the analog-injected playback path produces the correct monotonic ITD-to-direction mapping; it does not resolve whether the mapping remains reliable for actual sound sources in a room. Thus the central claim goes beyond the presented evidence, and the paper needs a quantitative real-world localization test.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript demonstrates a fully on-chip neuromorphic processing pipeline on the BrainScaleS-2 accelerated mixed-signal platform, applied to binaural sound localization. Two electret microphones feed analog, continuous-valued audio signals directly into on-chip neuron membranes, bypassing explicit ADC/DAC or event-based encoding. A spiking neural network implements a Jeffress-like delay-line coincidence detector that maps interaural time difference (ITD) to a spatial direction, and one of the chip's embedded microprocessors reads out the active coincidence neurons and generates a PWM/UART actuator signal. The authors report a measured end-to-end latency of 0.5 ms, a per-stage delay of 3.8 µs, and automated sweeps over ±149 µs of ITD showing a monotonic, approximately linear readout over 100 trials per ITD. Physical demonstrations with claps and a bouncing ping-pong ball are described qualitatively as aligning a servo-mounted owl figure toward the source. The central engineering novelty is the direct analog sensor injection and the fully on-chip closed-loop control, with the acceleration factor of BrainScaleS-2 used to make microsecond-scale ITDs accessible to spiking neurons.","tokens_in":8049,"tokens_out":5321,"duration_ms":65584,"significance":"If the real-world localization claim is substantiated quantitatively, this is a valuable hardware demonstration that advances near-sensor neuromorphic processing. The paper's strengths include the first reported direct injection of continuous sensor signals into the analog compute elements of BrainScaleS-2, the use of the embedded processors for actuator control in a closed loop, and a reproducible automated evaluation with 100 trials per ITD. The authors are also transparent about the synthetic-spike provenance of the rasterplot in Fig. 5B. The current evidence, however, supports the analog-injection and readout concept more strongly than the claimed ability to 'reliably predict the position of a sound ... in a room,' because the only quantitative localization data are obtained with software-imposed inter-channel delays rather than acoustic sources at known spatial positions. The significance for the neuromorphic hardware community hinges on closing this gap between the headline application and the quantified experiments.","major_comments":[{"comment":"The quantitative localization evaluation bypasses the acoustic scene. As stated in Section II-A, the automated sweep 'bypass[es] the microphones and play[s] back a recorded clap with an adjustable, software-set inter-channel delay.' This exercises the analog input path and the SNN readout, but not the claimed real-room scenario: variable source azimuth and distance, room reflections, reverberation, background noise, and angle-dependent spectral/amplitude cues are absent. The physical demonstrations in a room are anecdotal, with no measured angles, trial counts, or error statistics. The claim in the opening of Section III that the system 'can reliably predict the position of a sound ... in a room' therefore goes beyond the presented evidence. I request a quantitative real-world localization test with a sound source (e.g., loudspeaker or clapper) at several known azimuths, reporting per-an","section":"Section II-A and Section III, Fig. 5C/D"},{"comment":"The central evidence for the coincidence-detection mechanism is recorded with artificially injected spikes rather than the analog input path. The caption states: 'For technical reasons, panel (B) has been recorded with artificially injected spikes to the respective first neuron.' Since the paper's core contribution is direct analog injection, the rasterplots do not demonstrate that the analog-injected signals produce the same timing dynamics at the neuron membranes. If the analog path introduces latency, distortion, or altered effective time constants, the coincidence pattern shown may not hold for the real input. Although Fig. 5C/D show that the analog-injected playback yields a correct ITD-to-direction readout, they do not directly display the coincidence-neuron rasters or isolate the timing fidelity of the analog path. Please record the same rasterplot using the actual analog-injected","section":"Section III, Fig. 5B and caption"},{"comment":"The claimed 'theoretical spatial resolution of ≤1.5°' is derived from the measured per-stage delay (3.8±0.8 µs) combined with Woodworth's formula, but it is not validated against the actual readout statistics. The readout in Algorithm 1 averages the IDs of all coincidence neurons that fire within one 55 µs polling iteration, and the authors note that the temporal receptive field is broader than ideal. Fig. 5C/D show distribution widths of 2-4 units, yet the paper does not convert these widths into an angular error or confidence interval. The 'theoretical resolution' conflates the inter-stage delay spacing with the achievable localization accuracy. Please report the effective angular resolution derived from the measured output distributions (e.g., standard deviation or percentiles of detected direction per ITD) and discuss how it compares with the 1.5° figure. Also clarify whether Eq. (2)","section":"Section III, Table I and Eq. (2)"}],"minor_comments":[{"comment":"The term 'fully on-chip processing pipeline' is used in the abstract and Section IV, but the evaluation setup includes an external host computer, sound card, and analog preprocessing board. Please clarify that 'fully on-chip' refers to the neuromorphic processing chain (sensor input handling, SNN, readout, actuator signal generation) and not to the entire sensory acquisition and power path.","section":"Section II-A, Fig. 4"},{"comment":"Woodworth's formula uses θ in radians; please state this explicitly. Also, since the system uses two microphones without a head-like acoustic shadow, the direct use of the head-radius formula is an approximation; a short justification would help.","section":"Eq. (2)"},{"comment":"The sentence 'For either experiment, the coincidence neurons corresponding to the respective intersection of both delay chains become active' should read 'For each experiment...' since three stimuli are shown. Additionally, 'For technical reasons' is vague; a brief explanation of why artificial spikes were necessary would aid reproducibility.","section":"Section III, Fig. 5B"},{"comment":"The axis label 'detected direction [a.u.]' and the statement 'distributions' widths ... 2 to 4 units' are unclear. Specify whether the unit is a neuron ID index, an arbitrary analog unit, or an angle in degrees, and describe how the width was computed.","section":"Section III, Fig. 5C/D"},{"comment":"The phrase 'prefers early, high-confidence events' is imprecise: the algorithm only selects events that occur in the first polling iteration; it does not assign confidence weights, and the enumeration order introduces a bias that the authors state is not recognizable in practice. Consider rephrasing to avoid the implication of a principled confidence weighting.","section":"Algorithm 1"},{"comment":"The 0.5 ms latency ('first signal transient to digital output') is reported without a measurement method. Please specify how this was measured (e.g., oscilloscope trigger, event timestamps) and whether it includes the polling interval or only the SNN and readout delay.","section":"Section III, latency"}],"recommendation":"major_revision","confidential_remarks":"The engineering contribution is solid and likely of interest to the neuromorphic hardware community. The main gap is experimental, not conceptual: the automated evaluation validates the ITD-to-direction mapping for artificially delayed playback, but not the real-room acoustic localization claim. I would support acceptance after the authors add a quantitative test with physical sound sources at known angles and either reproduce the Fig. 5B raster with analog injection or provide an explicit validation of the equivalence. The spatial-resolution claim should also be derived from the measured output statistics rather than from the per-stage delay alone. These additions are within the scope of the current system and should be feasible with the existing demonstrator."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real hardware integration result—first direct injection of continuous analog audio into BrainScaleS-2's analog neuron circuits, plus on-chip processor generating servo PWM, so the whole sensor-to-actuator loop runs on the ASIC. That is the paper's contribution, and it is demonstrated convincingly as an engineering feat.\n\nThe Jeffress-style delay-line network is textbook, not the point. The point is that the analog path works end-to-end: they show monotonic ITD-to-direction readout over ±149 µs with 100 trials per ITD, outliers rare, and a 0.5 ms latency. They are also honest about the Fig. 5B rasterplots being recorded with artificially injected spikes, not the analog input. That is a limitation but not a fatal one, because Fig. 5C/D were recorded through the actual analog playback path, so the analog front-end is exercised. The reader's worry about analog-vs-spike equivalence is less load-bearing than the stress-test note suggests.\n\nThe real soft spot is the gap between the headline claim and the quantitative evidence. The only automated evaluation bypasses the microphones: they play back a recorded clap with software-set inter-channel delay. That tests the analog input stage and network, but not the acoustic scenario—room reflections, reverberation, source distance, ear-dependent amplitude cues. The live demos with claps and a ping-pong ball are anecdotal; no measured angles, trial counts, or error statistics. So \"reliably predict the position of a sound in a room\" is too strong for what is shown. That is an addressable flaw: add a simple real-room localization test with a known source azimuth and report errors.\n\nOne smaller point: the claimed 1.5° spatial resolution comes from the measured 3.8 µs per stage inserted into Woodworth's formula. That's a hardware parameter measurement, not a fit, but it's also not an end-to-end accuracy measurement, so it shouldn't be presented as the system's real-world resolution.\n\nOverall: someone working on near-sensor neuromorphic processing will get real value; the paper deserves a serious referee, but the authors should be asked to either temper the real-room claim or back it with quantitative data. I'd send it to review, not desk reject.","headline":"Genuine first demonstration of direct analog sensor injection into BrainScaleS-2 with on-chip actuator control, but the real-room localization claim outruns the quantitative evidence.","tokens_in":8459,"tokens_out":2183,"would_cite":false,"duration_ms":23865,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By feeding microphone voltages directly into the analog neuron circuits of an accelerated neuromorphic chip, this paper demonstrates a fully on-chip spiking-network pipeline that localizes transient sounds and drives a servo motor in real t","keywords":["neuromorphic hardware","analog signal injection","sound localization","interaural time difference","spiking neural network","BrainScaleS-2","near-sensor processing","real-time control"],"falsifier":"Inject a single calibrated click through each microphone channel and compare the first spike times of the input neurons against the same click delivered as artificial current pulses; if the analog-injected latency differs by more than a few microseconds, or if trial-to-trial jitter exceeds the roughly 3.8 µs delay per stage, then the coincidence detection shown in Fig. 5B is not representative of real inputs. Alternatively, place the sound source at a precisely known angle and check whether the detected direction matches Woodworth's formula within the claimed 2-4 units of the receptive field.","tokens_in":7634,"feed_emoji":"🔊","tokens_out":4096,"duration_ms":46033,"temperature":0.7,"pith_summary":"The paper shows that a mixed-signal neuromorphic chip can take continuous analog sensor voltages straight into its neuron circuits, bypassing the usual conversion to spikes or digital numbers. Using two microphones and a simplified Jeffress-style coincidence network, it turns the microsecond-scale time difference between the ears into a spatial code that predicts where a transient sound (a clap or a bouncing ball) came from. An on-chip microprocessor reads that code and drives a servo motor, so sensing, computation, and action all happen on one accelerator. The central motivation is efficiency: removing analog-to-digital and digital-to-analog conversions makes near-sensor processing faster and cheaper in energy. Because the hardware runs 1000 times faster than biology, its neurons have time constants that naturally resolve the small interaural delays that real-time millisecond-scale systems cannot.","feed_headline":"Chip finds a sound's direction in 0.5 ms, no A/D conversion","feed_subtitle":"Direct analog injection lets a spiking network hear, compute, and turn a servo toward a clap on one chip.","key_machinery":"The load-bearing mechanism is direct analog membrane injection: preprocessed microphone voltages are routed through an I/O pin onto the membrane of on-chip neurons, so continuous signals stimulate them without ADC/DAC or event conversion. The network is a Jeffress-style coincidence detector built from chains of LIF neurons with exponential synapses; the 1000x acceleration makes their 15 µs time constants natural delay elements for resolving interaural time differences of tens of microseconds. An embedded microprocessor polls the coincidence neurons' spike counters and converts their averaged IDs into a servo command.","core_discovery":"On the BrainScaleS-2 accelerated mixed-signal platform, the authors establish the first direct, continuous-valued sensory input into the analog compute units and use the chip's embedded processor for actuator control, creating a fully on-chip pipeline. A spiking neural network that implements a simplified Jeffress model—two counter-propagating chains of leaky-integrate-and-fire delay neurons feeding a row of coincidence detectors—maps interaural time difference to a spatial direction. With a 51 mm microphone spacing and a measured 3.8 µs per delay stage, the system resolves directions within roughly 1.5° in theory, shows near-linear response over ±149 µs of ITD, and moves a servo in real tim","pith_inferences":["The direct-injection principle should transfer to other continuous timing sensors (audio, vibration, ultrasound, bio-potentials), making the chip a general-purpose fast analog front end rather than a sound locator.","If on-chip learning rules were applied to the delay-chain weights or readout mapping, the system could self-calibrate microphone gain mismatches or adapt to changing geometry without host intervention.","A stronger test than the single-clap demo would be tracking a moving sound source; the 200 ms dead time and single-event readout may limit this, but the underlying coincidence code should support it with a faster readout scheme.","The paper characterizes coincidence dynamics using injected spikes; a direct comparison of analog-injected versus spike-injected responses is needed to separate network behavior from input-path artifacts."],"forward_implications":["A full sensory-motor loop can run on one chip without host involvement, reducing communication overhead and enabling low-power near-sensor robotics.","Eliminating A/D conversion removes a major energy and latency cost for tasks where sensors are inherently analog and timing is critical.","Because delays are implemented as programmable neuron chains, the same network can be retargeted to other ITD ranges, such as smaller microphone baselines for ultrasonic localization.","Network-level improvements like lateral inhibition or winner-take-all readout are directly implementable on the same hardware and would sharpen the coincidence resolution.","The 0.5 ms latency is dominated by the readout and polling design, so faster processor loops or parallel analog readout could reduce it further."],"fun_headline_variants":["Analog audio in, servo out: one chip does real-time localization","No A/D conversion: chip hears and aims in real time","Accelerated neuromorphic chip processes sound directly to servo control","Direct analog input: spiking network locates sounds, moves servo","On-chip sound localization: analog injection meets 1000x acceleration"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the analog-injected audio signal produces the same timing at the neuron membranes as the synthetic spikes used to record the network's coincidence behavior; if the analog conditioning path distorts waveshape or adds uncontrolled delay, the measured 3.8 µs per stage and the coincidence rasters may not hold for real microphone input.","fun_headline_variants_meta":{"raw":{"variants":["Analog audio in, servo out: one chip does real-time localization","No A/D conversion: chip hears and aims in real time","Accelerated neuromorphic chip processes sound directly to servo control","Direct analog input: spiking network locates sounds, moves servo","On-chip sound localization: analog injection meets 1000x acceleration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000659,"raw_usage":{"total_tokens":2855,"prompt_tokens":754,"completion_tokens":2101,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":2012}},"tokens_in":498,"tokens_out":2101,"duration_ms":15719,"temperature":1.0,"reasoning_tokens":2012,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T04:30:56.485216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject a single calibrated click through each microphone channel and compare the first spike times of the input neurons against the same click delivered as artificial current pulses; if the analog-injected latency differs by more than a few microseconds, or if trial-to-trial jitter exceeds the roughly 3.8 µs delay per stage, then the coincidence detection shown in Fig. 5B is not representative of real inputs. Alternatively, place the sound source at a precisely known angle and check whether the detected direction matches Woodworth's formula within the claimed 2-4 units of the receptive field.","supporting_citations":[],"review_version":1}