REVIEW 3 major objections 3 minor 1 cited by
Training nonlinear optical neural networks with Scattering Backpropagation
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Scattering Backpropagation trains nonlinear optical neural networks using only two scattering experiments, with no mathematical model of the nonlinearity required.
desk verdict Plausible, well-framed idea that two scattering experiments can yield model-free gradients for nonlinear optical networks; worth a real referee, but the abstract alone proves nothing and the prior-art claim needs scrutiny. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the two-scattering-experiment procedure underlying Scattering Backpropagation. The network is treated as a scattering medium, and two carefully chosen scattering measurements are combined to extract approximations of all parameter gradients. The combination exploits reciprocity of the physical system to relate the two measured field distributions, which is why no model of the nonlinearity is required; the deviation from reciprocity directly sets the gradient-estimation error.
What would settle it
Take a nonlinear optical network with a tunable nonreciprocal element and compare the gradients from Scattering Backpropagation with exact gradients from a full simulator. The paper's claim predicts that the estimation error grows monotonically with the degree of nonreciprocity; if the error stays flat or shrinks when reciprocity is broken, the central mechanism is falsified.
Extended reading notes
Core claim
The paper's central claim is that one can train the most general class of nonlinear optical neural networks by measuring approximate gradients from only two scattering experiments. No analytical or numerical model of the nonlinearity is needed; the physical system itself supplies the gradient information. The approximation error is controlled by the system's deviation from reciprocity, so the method is most trustworthy for reciprocal networks. The authors validate the method on XOR and MNIST, showing that it reaches usable accuracy, and they point to existing scalable platforms—optics, microwaves, and electrical circuits—as natural targets.
Load-bearing premise
The whole gradient estimate depends on the optical network behaving nearly reciprocally; if reciprocity is strongly violated, the two scattering experiments give biased gradients and training would fail.
Editorial extensions
If this is right
- Training nonlinear optical networks no longer requires a mathematical model of the nonlinearity, removing a major bottleneck for hardware implementation.
- Only two scattering experiments are needed to extract all gradient approximations, making the method efficient and independent of the number of parameters.
- The method succeeds on standard benchmarks XOR and MNIST, suggesting it is generic enough for practical tasks.
- Because it uses only scattering measurements, the method can transfer to other platforms such as microwave circuits and electrical circuits.
- Gradient precision is predictable: it is tied to reciprocity, so designing reciprocal networks keeps training reliable.
Reading between the lines
- By extension, this could enable in-situ training of photonic accelerators on noisy or fabrication-imperfect hardware, since the hardware itself provides the gradients instead of a simulation.
- The reciprocity-dependent error suggests an experimental test: deliberately break reciprocity in a test network and watch training degrade; a monotone relationship would confirm the mechanism.
- Similar two-measurement schemes might be applicable to any reciprocal wave-based physical learning machine, such as acoustic or mechanical networks, where exact gradients are hard to obtain.
- If the method scales, energy savings would compound: no backpropagation computation on a digital computer, only two optical experiments per update.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript proposes Scattering Backpropagation, a training method for nonlinear optical neural networks that estimates gradients from two scattering experiments, without requiring a mathematical model of the physical nonlinearity. The authors state that gradient-estimation precision depends on the deviation from reciprocity and demonstrate the method on XOR and MNIST benchmarks. This review is based solely on the abstract; the full text was not available.
Significance. If the central claim holds, the method would fill a notable gap: currently there is no efficient, generic, physics-based training algorithm for the broad class of nonlinear optical systems. The promise of extracting all gradient approximations from only two scattering experiments, without a model of the nonlinearity, is attractive for scalable optical neuromorphic hardware. The claimed extension to microwave and electrical-circuit platforms increases the potential impact. However, the abstract alone does not allow verification of the derivation, the error bounds, or the experimental validity, so the significance cannot yet be fully assessed.
major comments (3)
- [Abstract (central claim)] The central claim—'only involves two scattering experiments to extract all gradient approximations'—is not derived or qualified. The abstract does not specify what is measured in each experiment, how the gradient estimator is constructed, or whether 'all gradient approximations' means all parameters simultaneously or per layer. A derivation with explicit assumptions (e.g., weak nonlinearity, reciprocity conditions) and an error bound is needed before this claim can be evaluated.
- [Abstract (reciprocity dependence)] The statement that 'estimation precision depends on the deviation from reciprocity' is the main stated limitation, but the abstract gives no quantitative relation. If the bias grows linearly with the non-reciprocal response, many practical systems (e.g., magneto-optical) may fail to train; if the dependence is higher-order, the useful regime may be broad. The manuscript must provide a bound or experimental characterization of this dependence.
- [Abstract (benchmarks)] The XOR and MNIST results are reported without error bars, comparison to baselines (e.g., exact backpropagation, physics-aware training, other hardware-in-the-loop methods), or number of trials. Since 'successfully apply' is offered as evidence for the method, these experimental details are necessary to judge whether the results support the claim.
minor comments (3)
- [Abstract] 'Two scattering experiments' is ambiguous; clarify whether these are forward/adjoint-type measurements and whether the number remains exactly two for networks of arbitrary depth and width or is per layer.
- [Abstract] 'All gradient approximations' could be read as exact gradients; the qualified term 'approximate gradients' should be used consistently and explicitly.
- [Abstract] The sentence on applicability to 'optics, microwave, and also extends to other physical platforms such as electrical circuits' would benefit from a one-sentence explanation of why the same scattering formalism applies to these platforms.
Circularity Check
No circularity identifiable from abstract; derivation chain not available for inspection.
full rationale
The review is restricted to the abstract of arXiv:2508.11750, as no full text was provided. The abstract claims that Scattering Backpropagation extracts approximate gradients for nonlinear optical neural networks using two scattering experiments and without requiring a mathematical model of the nonlinearity, with estimation precision depending on deviation from reciprocity. This is a stated methodological claim plus an acknowledged limitation. Nothing in the abstract indicates that a fitted parameter is renamed as a prediction, that a central quantity is defined in terms of the target, that a load-bearing premise rests only on self-citation, or that a known result is repackaged. The reciprocity-dependence caveat is a limitation on validity, not a circular step. Per the hard rules, circularity may only be flagged when specific quoted equations or constructions exhibit the reduction; no such evidence exists in the available text. An honest non-finding is therefore appropriate: the claimed approach is not shown to be self-contained from the abstract, but no circularity can be established without the derivation chain.
Assumptions & free parameters
assumptions (2)
- domain assumption The nonlinear optical network's behavior can be characterized through scattering measurements.
- domain assumption The gradient estimation error is bounded by the deviation from reciprocity.
Cite this review
Pith. "Pith review of Training nonlinear optical neural networks with Scattering Backpropagation." pith.science (2026). https://pith.science/paper/ZQ55W3AD
@misc{pith2026250811750,
author = {Pith},
title = {Pith review of: Training nonlinear optical neural networks with Scattering Backpropagation},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZQ55W3AD}},
note = {Machine review of arXiv:2508.11750}
}
read the original abstract
As deep learning applications continue to deploy increasingly large artificial neural networks, the associated high energy demands are creating a need for alternative neuromorphic approaches. Optics and photonics are particularly compelling platforms as they offer high speeds and energy efficiency. Neuromorphic systems based on nonlinear optics promise high expressivity with a minimal number of parameters. However, so far, there is no efficient and generic physics-based training method allowing us to extract gradients for the most general class of nonlinear optical systems. In this work, we present Scattering Backpropagation, an efficient method for experimentally measuring approximated gradients for nonlinear optical neural networks. Remarkably, our approach does not require a mathematical model of the physical nonlinearity, and only involves two scattering experiments to extract all gradient approximations. The estimation precision depends on the deviation from reciprocity. We successfully apply our method to well-known benchmarks such as XOR and MNIST. Scattering Backpropagation is widely applicable to existing state-of-the-art, scalable platforms, such as optics, microwave, and also extends to other physical platforms such as electrical circuits.
Forward citations
Cited by 1 Pith paper
-
Equilibrium Propagation for Non-Conservative Systems
A modified Equilibrium Propagation with an antisymmetric-Jacobian correction computes exact cost gradients for non-conservative neural dynamics.
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.