Pith. sign in

REVIEW 3 major objections 4 minor 103 references

The paper claims that a strictly local, excitatory Hebbian learning rule, evaluated as a synaptic resource-allocation mechanism, produces representations with lower task-information cost than sparse backpropagation and target propagation at

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 21:33 UTC pith:ROKGMIR6

load-bearing objection New empirical comparison with a plausible confound in the CTI metric; the resource-allocation story is provisional until bound-tightness is checked, but the paper deserves a serious referee. the 3 major comments →

arxiv 2607.16027 v1 pith:ROKGMIR6 submitted 2026-07-17 cs.LG cs.NE

Constrained Hebbian Learning Supports Efficient Representational Allocation under Structural Constraints

classification cs.LG cs.NE
keywords Hebbian learningsynaptic resource allocationvariational information bottlenecktask-information costsparse neural networksnonnegative weightsauditory-visual learninglocal plasticity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper claims that a strictly local, excitatory Hebbian learning rule — the Oja-style update of Eq. (5) with nonnegative weights and competitive normalization — allocates representational resources more efficiently than backpropagation (BP) and difference target propagation (DDTP) when networks are constrained by sparsity and Dale's law. Efficiency is measured by Task-Information Cost, CTI = I(Z;H)/I(Z;Y), the amount of input information retained in a VIB latent code per unit of task-relevant information, estimated post hoc on frozen representations. On three audiovisual benchmarks (AVE, Kinetics-Sounds, VGGSound100) with fixed upstream embeddings, Hebbian-trained MLPs reach substantially lower CTI than sparse BP and DDTP (geometric ratios 5.4 and 6.7) at comparable or slightly lower Top-1 accuracy, and their representations degrade less under magnitude pruning. If the metric is trustworthy, the result supports viewing Hebbian plasticity as a local mechanism for synaptic resource allocation under metabolic constraints, not as a general accuracy-maximizing strategy; the authors themselves stress CTI is a comparative proxy, not a direct energy measure, and limit the claim to shallow and five-layer compressed settings.

Core claim

Under matched sparsity (~10% connectivity), nonnegativity, and identical MLP capacity, a local excitatory Hebbian update yields post hoc VIB codes with Task-Information Cost (CTI) roughly 9–23 across datasets and architectures, versus 25–344 for sparse BP and 46–210 for DDTP: as much task-relevant information, far less retained input information. Paired log-ratio tests are significant after Holm correction; accuracy is close but not uniformly higher. The authors frame this as a cost-performance trade-off, note that shallow nonnegative BP matches the low-CTI regime but fails deep, and report a ten-hidden-layer Hebbian control that does not stabilize. CTI is explicitly a comparative proxy, not

What carries the argument

The carrying mechanism is the constrained Hebbian update of Eq. (5), an Oja-style neural-PCA rule whose subtractive term −η z_j Σ_k z_k w_ik decorrelates postsynaptic activity. Applied with nonnegative weights, min–max rescaling of activations to [0,1], and per-layer z-score normalization, it yields sparse, Dale's-law-compliant, decorrelated hidden codes. The post hoc Variational Information Bottleneck (VIB) module — a diagonal-Gaussian stochastic encoder trained with a KL-to-prior compression penalty and a linear decoder — converts each frozen representation into a latent Z; CTI = I(Z;H)/I(Z;Y), using the KL proxy for I(Z;H) and a variational lower bound for I(Z;Y), turns that code into a s

Load-bearing premise

The load-bearing premise, spelled out in Secs. 3.2 and 3.6.1, is that CTI = I(Z;H)/I(Z;Y) — an upper-bound KL proxy divided by a variational lower bound — is a valid and comparable measure of representational cost across learning rules; if the Hebbian activations' zero-centered bounded shape makes the Gaussian VIB fit artificially better, the reported CTI advantage is a measurement artifact rather than a property of the rule.

What would settle it

Take the three frozen representation sets (Hebbian, sparse BP, DDTP) for one dataset and architecture, and compute I(Z;H) not with the KL-to-N(0,I) proxy but with the aggregated-posterior KL or a nonparametric estimator (e.g., k-NN) on samples from the VIB encoder, then recompute CTI at matched β and accuracy. If the Hebbian CTI advantage over BP/DDTP vanishes or reverses, the paper's resource-allocation conclusion is an estimation artifact.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A purely local Hebbian rule operating under cortical-like sparsity can keep task-relevant information without global error signals, which would mean brains can assemble efficient associative codes without backpropagation.
  • Representational cost can be compared across learning rules independently of raw accuracy: two models with the same Top-1 accuracy can differ by 4–5× in CTI, so accuracy alone is insufficient to evaluate biologically constrained learners.
  • Hebbian-trained representations tolerate post hoc pruning down to ~10% connectivity with little accuracy loss, while BP and DDTP degrade earlier, implying the learned connectivity is already close to its functional support.
  • Bimodal audiovisual inputs can raise accuracy without proportionally raising CTI, suggesting that cross-modal integration can be representationally inexpensive under this proxy.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The Gaussian VIB posterior may fit the Hebbian activations' zero-centered, bounded [0,1] shape better than it fits BP/DDTP activations, so the CTI gap could shrink or vanish under a nonparametric estimator of I(Z;H); checking this is the most direct test of the resource-allocation interpretation.
  • If CTI tracks synaptic maintenance cost as the paper's thermodynamic analogy suggests, Hebbian-trained networks should show measurably lower energy use than BP/DDTP at matched sparsity on neuromorphic hardware; this is a concrete hardware prediction the paper does not test.
  • The ten-hidden-layer failure suggests the bottleneck readout, not the Hebbian rule, may be removing task-relevant information at depth; adding an explicit inhibitory population (≈20%) or feedback connections, as the paper sketches, could recover depth while preserving low CTI.
  • Because the inputs are fixed MAE embeddings, the result is about downstream associative plasticity, not end-to-end feature discovery; applying the same assay to raw audiovisual streams would test whether the resource-allocation advantage survives upstream learning.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper tests whether a strictly local, excitatory, competitive Hebbian rule (a variant of Oja's PCA rule with nonnegativity and normalization) allocates representational resources more efficiently than backpropagation (BP) and Dense Difference Target Propagation (DDTP) under matched sparsity and architectural constraints. Using fixed VideoMAE/AudioMAE embeddings from AVE, Kinetics-Sounds, and VGGSound100, the authors train shallow and deep MLPs with each rule, freeze the encoder, and fit a post hoc Variational Information Bottleneck (VIB) module. The central metric is Task-Information Cost, CTI = I(Z;H)/I(Z;Y), estimated from VIB bounds. The main empirical claim is that Hebbian networks achieve lower CTI than sparse BP and DDTP in compressed comparisons, with accuracy remaining comparable in several settings, and that this supports interpreting Hebbian plasticity as a resource-allocation mechanism rather than a general accuracy-maximizing strategy.

Significance. If the central claim holds, the paper would provide a valuable demonstration that a biologically local Hebbian rule can shift the information-efficiency trade-off in a controlled assay, with implications for synaptic resource allocation theories. Strengths include the two-stage protocol that separates representation learning from the information-theoretic readout, matched sparsity and architecture across rules, paired seed-level statistical testing with Holm correction, and a set of ablations (MNIST, GHA, classical Oja, PCA-to-readout, depth-scaling). The manuscript is also unusually candid about its limitations, including the instability of deep nonnegative BP and the 10-layer Hebbian control. However, the central quantitative claim rests on a CTI estimate whose numerator and denominator are variational bounds of opposite direction; if bound tightness differs systematically across learning rules, the reported Hebbian advantage could be a measurement artifact rather than a property of the learning rule. The significance is therefore contingent on a matched-estimator or bound-tightness check.

major comments (3)
  1. [§3.2 and §3.6.1, Eq. (9)] CTI is defined as I(Z;H)/I(Z;Y), but the implementation estimates I(Z;H) by E_h KL[q(z|h)||r(z)] = I(Z;H) + KL[q(z)||r(z)] (Eq. 4, an upper bound) and I(Z;Y) by a variational lower bound. The ratio of an upper bound to a lower bound is not a bound on the true ratio, and the two gaps can differ across learning rules. The paper's own preprocessing creates a concrete mechanism: Hebbian and nonnegative-BP activations are Z-score normalized and min-max rescaled to [0,1] (Secs. 3.1, 3.5), while BP/DDTP use tanh activations without rescaling. Since the VIB encoder is a single linear layer into diagonal-Gaussian parameters, the scale and distribution of h directly affect how tightly a diagonal Gaussian posterior can fit. Thus the large Hebbian advantage in Table 1 (geometric CTI ratios 5.37 and 6.73) may reflect smaller bound gaps for Hebbian representations rather than genuinely lower informati
  2. [§4.1.1, Fig. 2 and Table 1] The main tabular comparison is selected at β=10^-2, and the text states that higher-β operating points for DDTP and BP were excluded because they show a rapid decline in performance. Since CTI generally decreases as β increases, excluding the higher-β points of the reference methods removes exactly the operating points where BP/DDTP might achieve lower CTI. To support the claim that Hebbian learning 'achieves lower CTI than sparse BP and DDTP,' the authors should report complete β trajectories and compare operating points on a comparable basis (e.g., matched I(Z;Y) or matched Top-1 accuracy), not only at a fixed β chosen post hoc. Without this, the reported advantage may be an artifact of the selected operating point.
  3. [§3.6.1, statistical analysis] The paired Wilcoxon analysis (N=30, pHolm=3.73e-9) is correctly applied to log-transformed CTI ratios and supports the claim that the Table 1 values differ in the matched conditions. However, the analysis inherits the validity of the CTI estimator. If the bound-tightness concern in the first major comment is not addressed, the statistical test only shows that the estimated CTI values differ, not that the true information costs differ. The authors should either defend the comparability of the bounds more rigorously or report the statistical test on a corrected estimator. This is not a call to remove the statistics, but to ensure the quantity being tested is the quantity of scientific interest.
minor comments (4)
  1. [§3.1, Eq. (5)] The activation function φ(a_j) is defined as a min-max rescaling over 'the corresponding hidden layer,' but it is not stated whether the min/max are computed per batch, per dataset, or over a running statistic. Since the VIB input distribution depends on this, please clarify the exact normalization procedure in the main text.
  2. [Table 1] The row 'BP (nonneg.) Deep*' contains only em-dashes, and the footnote says dense connectivity only for the first BP row. This is confusing. If deep nonnegative BP was not trained on AVE/Kinetics-Sounds and only a single run exists on VGGSound100, state this explicitly in a table note rather than using an empty row.
  3. [§4.1.4, Table 4] The 10-hidden-layer depth-scaling control reports an approximate Top-1 range of 7–10% with no multi-seed estimate. Since this is interpreted as a limitation, it would be helpful to state the chance level for the 100-class VGGSound100 split (1%) so readers can assess how close to chance that range is.
  4. [Global] There are several typographical issues: 'T raining Algorithm' in Section 3.5.1, 'A VE' with a space throughout, and inconsistent use of 'nonneg.' vs. 'nonnegativity-constrained' in tables. These do not affect the science but should be cleaned up.

Circularity Check

0 steps flagged

No load-bearing circularity; main derivation is empirically self-contained, with minor self-citations and an estimator-comparability caveat.

full rationale

The central claim—that Hebbian representations occupy a lower-CTI regime than sparse BP and DDTP—is an empirical comparison, not a derivation that reduces to its inputs. Hebbian weights are learned in Phase 1 from Eq. (5) without access to the VIB readout; CTI is then measured post hoc (Sec. 3.2, Eq. (6); Sec. 3.6.1, Eq. (9)). The rule is specified in the paper (Eq. (5), Algorithm 1), so citations [22,23] to the authors' earlier work are provenance rather than load-bearing evidence. External controls (MNIST, PCA-to-readout, nonnegativity-constrained BP, ablation without normalization/rescaling) provide independent checks. The main caveat is not circularity: CTI is a ratio of a variational upper bound for I(Z;H) to a variational lower bound for I(Z;Y), so cross-rule differences in bound tightness—possibly tied to the [0,1]/Z-score preprocessing used only for nonnegative networks—could bias the comparison. This is a measurement-validity risk requiring a matched-estimator check, but it does not make any equation equivalent to another by construction. Minor self-citations exist but are not used to force the conclusion, so the score is low.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

The central claim rests on the VIB estimator and the β selection more than on a derivation. Free parameters are mostly standard hyperparameters, but their per-rule tuning and post hoc β choice directly affect the reported CTI gaps. The axioms are domain assumptions about the information-theoretic proxies; none are machine-checked.

free parameters (5)
  • β (VIB Lagrange multiplier) = 10^-2 (main comparison); swept as 0, 10^-3, 10^-2, 10^-1
    Controls the compression-relevance trade-off; the main tabular result uses β=10^-2, selected post hoc as the compressed regime of interest.
  • Per-rule learning rates = η=1e-3 BP, 5e-5 Hebbian, 1e-2 DDTP
    Chosen via a brief randomized search on a validation subset; learning rates are not matched across rules, so accuracy/cost differences could partly reflect optimization tuning.
  • Latent dimension K = 256
    Taken from [16] after 'initial experiments' confirmed it; affects VIB estimates of I(Z;H) and I(Z;Y).
  • Sparsity target = ~90%
    Imposed post hoc on all networks; chosen from cortical connectivity estimates, but pruning threshold affects accuracy and CTI.
  • VIB scale offset = 5.0 in softplus
    Initialization choice for the stochastic encoder; inherited from [16].
axioms (6)
  • standard math Oja's rule converges to principal components when inputs are mean-centered.
    Invoked in Section 3.5 to justify zero-centering before Hebbian updates (ref [57]).
  • domain assumption The expected KL to the prior approximates I(Z;H) up to KL(q(z)||r(z)).
    Section 3.2, Eq. (4); posterior is diagonal Gaussian with fixed prior, so this is an upper bound, not an exact mutual information.
  • domain assumption The variational lower bound H(Y) - E[-log q(y|z)] approximates I(Z;Y).
    Section 3.2; the bound's tightness may vary across learning rules, which is load-bearing for CTI comparisons.
  • domain assumption Reduced representational complexity (lower entropy/information) maps to reduced metabolic cost.
    Sections 1 and 2.1; the paper explicitly states CTI is a comparative proxy, not a direct energy measurement.
  • domain assumption Fixed VideoMAE/AudioMAE embeddings stand in for fixed early sensory processing.
    Section 3.3; the audiovisual results are conditional on this abstraction, and the MNIST control only partially addresses raw-input behavior.
  • ad hoc to paper Normalization and activation rescaling do not confound the CTI comparison.
    Nonnegativity-constrained networks receive z-score normalization and [0,1] rescaling, while unconstrained BP/DDTP use tanh and different input scaling; the shallow nonnegative-BP control partially addresses this, but the deep comparison lacks a matched nonnegative baseline.

pith-pipeline@v1.3.0-alltime-deepseek · 29403 in / 13620 out tokens · 125625 ms · 2026-08-01T21:33:19.219060+00:00 · methodology

0 comments
read the original abstract

Introduction: Biological systems face anatomical and metabolic constraints, including costly synaptic maintenance and limited connectivity. These constraints favor neural codes that compress behaviorally relevant information into low-redundancy patterns. We test whether an excitatory competitive Hebbian rule can support synaptic resource allocation under such constraints and whether the resulting representations occupy a more favorable cost-performance regime than reference learning rules. Methods: Representational cost is quantified using mutual-information-based measures derived from the Variational Information Bottleneck. Experiments use fixed audiovisual embeddings from three audiovisual benchmarks (AVE, Kinetics-Sounds, VGGSound100) to isolate downstream associative plasticity. Hebbian learning is compared with Dense Difference Target Propagation (DDTP) and backpropagation (BP) under matched sparsity and architectural constraints. Results: Hebbian learning achieves lower task-information cost (CTI) than sparse BP and DDTP in the main compressed comparisons, while reaching CTI values comparable to shallow BP with nonnegative weights. Rather than uniformly improving classification performance, Hebbian learning shifts the trade-off between task-relevant information and representational cost, yielding lower CTI at comparable functional performance in several settings. Discussion: The results indicate a cost-performance trade-off rather than uniform accuracy gains. For a given level of task-relevant information, Hebbian representations retain less input information while preserving functional performance, although accuracy is slightly reduced on some datasets. These findings support interpreting Hebbian learning as a mechanism for synaptic resource allocation rather than as a general strategy for maximizing audiovisual classification accuracy.

Figures

Figures reproduced from arXiv: 2607.16027 by Andreas Knoblauch, Florian R\"ohrbein, Patrick Inoue.

Figure 1
Figure 1. Figure 1: Two-stage training protocol. Phase 1 learns deterministic repre [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Task-relevant information I(Z; Y ) versus representational cost I(Z; H) for networks trained with DDTP, BP, and the Hebbian rule under study. Points are connected within each learning-rule trajectory according to increasing bottleneck strength. For each trajectory, the points are ordered from top to bottom as β = 0, 10−3 , 10−2 , 10−1 , where β = 0 denotes the setting without a VIB compression penalty. The… view at source ↗
Figure 3
Figure 3. Figure 3: Top-1 accuracy after the Phase-2 VIB readout versus layer-wise [PITH_FULL_IMAGE:figures/full_fig_p025_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

103 extracted references · 28 canonical work pages · 3 internal anchors

  1. [1]

    Information and Effi- ciency in the Nervous System–A Synthesis,

    B. Sengupta, M. Stemmler, and K. J. Friston, “Information and Effi- ciency in the Nervous System–A Synthesis,”PLoS Comput. Biol., vol. 9, p. e1003157, Jul. 2013, doi: 10.1371/journal.pcbi.1003157

  2. [2]

    Structural Plasticity, Effectual Connectivity and Memory in Cortex,

    A. Knoblauch and F. T. Sommer, “Structural Plasticity, Effectual Connectivity and Memory in Cortex,”Front. Neuroanat., vol. 10, p. 63, 2016, doi: 10.3389/fnana.2016.00063. Available:https://www. frontiersin.org/articles/10.3389/fnana.2016.00063/full

  3. [3]

    An Energy Budget for Signaling in the Grey Matter of the Brain,

    D. Attwell and S. Laughlin, “An Energy Budget for Signaling in the Grey Matter of the Brain,”J. Cereb. Blood Flow Metab., vol. 21, pp. 1133–1145, Oct. 2001, doi: 10.1097/00004647-200110000-00001

  4. [4]

    The Cost of Cortical Computation,

    P. Lennie, “The Cost of Cortical Computation,”Curr. Biol., vol. 13, pp. 493–497, Apr. 2003, doi: 10.1016/S0960-9822(03)00135-0

  5. [5]

    Interpreting Oxygenation-Based Neuroimaging Sig- nals: The Importance and the Challenge of Understanding Brain Oxygen Metabolism,

    R. B. Buxton, “Interpreting Oxygenation-Based Neuroimaging Sig- nals: The Importance and the Challenge of Understanding Brain Oxygen Metabolism,”Front. Neuroenergetics, vol. 2, p. 8, 2010, doi: 10.3389/fnene.2010.00008

  6. [6]

    The Thermodynamics of Thinking: Connections Be- tween Neural Activity, Energy Metabolism and Blood Flow,

    R. B. Buxton, “The Thermodynamics of Thinking: Connections Be- tween Neural Activity, Energy Metabolism and Blood Flow,”Phi- los. Trans. R. Soc. B, vol. 376, no. 1815, p. 20190624, 2021, doi: 10.1098/rstb.2019.0624

  7. [7]

    Receptive Fields and Functional Ar- chitecture of Monkey Striate Cortex,

    D. H. Hubel and T. N. Wiesel, “Receptive Fields and Functional Ar- chitecture of Monkey Striate Cortex,”J. Physiol., vol. 195, no. 1, pp. 215–243, 1968, doi: 10.1113/jphysiol.1968.sp008455

  8. [8]

    Edge Co-Occurrence in Natural Images Predicts Contour Group- ing Performance,

    W. S. Geisler, J. S. Perry, B. J. Super, and D. P. Gallogly, “Edge Co-Occurrence in Natural Images Predicts Contour Group- ing Performance,”Vis. Res., vol. 41, no. 6, pp. 711–724, 2001, doi: 10.1016/S0042-6989(00)00263-2. 34

  9. [10]

    Emergence of Simple-Cell Recep- tive Field Properties by Learning a Sparse Code for Natural Images,

    B. A. Olshausen and D. J. Field, “Emergence of Simple-Cell Recep- tive Field Properties by Learning a Sparse Code for Natural Images,” Nature, vol. 381, pp. 607–609, 1996, doi: 10.1038/381607a0

  10. [11]

    Memory Capacities for Synaptic and Structural Plasticity,

    A. Knoblauch, G. Palm, and F. T. Sommer, “Memory Capacities for Synaptic and Structural Plasticity,”Neural Comput., vol. 22, no. 2, pp. 289–341, 2010, doi: 10.1162/neco.2009.08-07-588

  11. [12]

    Die Thermodynamik chemischer Vorg¨ ange,

    H. von Helmholtz, “Die Thermodynamik chemischer Vorg¨ ange,” inSitzungsberichte der K¨ oniglich Preußischen Akademie der Wis- senschaften zu Berlin, pp. 22–39, 1882

  12. [13]

    Energy Efficiency and Coding of Neu- ral Network,

    S. Li, C. Yan, and Y. Liu, “Energy Efficiency and Coding of Neu- ral Network,”Front. Neurosci., vol. 16, p. 1089373, Jan. 2023, doi: 10.3389/fnins.2022.1089373

  13. [14]

    The Free-Energy Principle: A Unified Brain The- ory?,

    K. J. Friston, “The Free-Energy Principle: A Unified Brain The- ory?,”Nat. Rev. Neurosci., vol. 11, no. 2, pp. 127–138, 2010, doi: 10.1038/nrn2787

  14. [15]

    A Free Energy Principle for Biological Sys- tems,

    K. J. Friston, “A Free Energy Principle for Biological Sys- tems,”Entropy, vol. 14, no. 11, pp. 2100–2121, 2012, doi: 10.3390/e14112100. Available:https://www.ncbi.nlm.nih.gov/ pmc/articles/PMC3510653/

  15. [16]

    Deep Varia- tional Information Bottleneck,

    A. A. Alemi, I. Fischer, J. V. Dillon, and K. Murphy, “Deep Varia- tional Information Bottleneck,” inProc. Int. Conf. Learn. Represent. (ICLR), OpenReview.net, 2017. Available:https://openreview. net/forum?id=HyxQzBceg

  16. [17]

    Impact of Structural Plasticity on Memory Formation and Decline,

    A. Knoblauch, “Impact of Structural Plasticity on Memory Formation and Decline,” inRewiring the Brain: A Computational Approach to Structural Plasticity in the Adult Brain, A. van Ooyen and M. Butz, Eds. Elsevier/Academic Press, 2017, pp. 361–386, doi: 10.1016/B978- 0-12-803784-3.00017-2

  17. [18]

    The Lottery Ticket Hypothesis: Find- ing Sparse, Trainable Neural Networks,

    J. Frankle and M. Carbin, “The Lottery Ticket Hypothesis: Find- ing Sparse, Trainable Neural Networks,” inProc. Int. Conf. Learn. Represent. (ICLR), 2019. Available:https://arxiv.org/abs/1803. 03635. 35

  18. [19]

    Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural Networks,

    X. Chang, Y. Li, S. Oymak, and C. Thrampoulidis, “Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural Networks,”Proc. AAAI Conf. Artif. Intell., vol. 35, pp. 6974–6983, May 2021, doi: 10.1609/aaai.v35i8.16859

  19. [20]

    Learning Both Weights and Connections for Efficient Neural Networks,

    S. Han, J. Pool, J. Tran, and W. Dally, “Learning Both Weights and Connections for Efficient Neural Networks,” inAdv. Neural Inf. Pro- cess. Syst., vol. 28, 2015

  20. [21]

    Pruning Filters for Efficient ConvNets,

    H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning Filters for Efficient ConvNets,” inProc. Int. Conf. Learn. Represent. (ICLR), 2017. Available:https://arxiv.org/abs/1608.08710

  21. [22]

    Guiding Sparse Neu- ral Networks with Neurobiological Principles to Elicit Biologically Plausible Representations,

    P. Inoue, F. R¨ ohrbein, and A. Knoblauch, “Guiding Sparse Neu- ral Networks with Neurobiological Principles to Elicit Biologically Plausible Representations,” arXiv preprint arXiv:2603.03234, 2026, doi: 10.48550/arXiv.2603.03234. Available:https://arxiv.org/ abs/2603.03234

  22. [23]

    Energy-Efficient Informa- tion Representation in MNIST Classification Using Biologically In- spired Learning,

    P. Inoue, F. R¨ ohrbein, and A. Knoblauch, “Energy-Efficient Informa- tion Representation in MNIST Classification Using Biologically In- spired Learning,” inProceedings of the 10th bwHPC Symposium, M. Janczyk, D. von Suchodoletz, B. Wiebelt, and M. Frank, Eds. KIT Scientific Publishing, 2026, doi: 10.58895/ksp/1000169488-2

  23. [24]

    Weight Perturbation and Competitive Hebbian Plasticity for Training Sparse Excitatory Neural Networks,

    P. Stricker, F. R¨ ohrbein, and A. Knoblauch, “Weight Perturbation and Competitive Hebbian Plasticity for Training Sparse Excitatory Neural Networks,” inProc. IEEE Int. Joint Conf. Neural Netw. (IJCNN), Yokohama, Japan, 2024, doi: 10.1109/IJCNN60899.2024.10650478

  24. [25]

    Solving the Problem of Negative Synaptic Weights in Cortical Models,

    C. Parisien, C. Anderson, and C. Eliasmith, “Solving the Problem of Negative Synaptic Weights in Cortical Models,”Neural Comput., vol. 20, pp. 1473–1494, Jun. 2008, doi: 10.1162/neco.2008.07-06-295

  25. [26]

    A Theoretical Framework for Target Propagation,

    A. Meulemans, F. Carzaniga, J. Suykens, J. Sacramento, and B. F. Grewe, “A Theoretical Framework for Target Propagation,” inAdv. Neural Inf. Process. Syst., vol. 33, pp. 20024–20036, 2020. Available: https://proceedings.neurips.cc/paper_files/paper/2020/ file/e7a425c6ece20cbc9056f98699b53c6f-Paper.pdf

  26. [27]

    Audio-Visual Event Localization in Unconstrained Videos,

    Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-Visual Event Localization in Unconstrained Videos,” inProc. Eur. Conf. Comput. Vis. (ECCV), pp. 252–268, 2018, doi: 10.1007/978-3-030-01216-8 16

  27. [28]

    Look, Listen and Learn,

    R. Arandjelovi´ c and A. Zisserman, “Look, Listen and Learn,” in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), pp. 609–617, 2017, doi: 10.1109/ICCV.2017.73. 36

  28. [29]

    VGGSound: A Large-Scale Audio-Visual Dataset,

    H. Chen, W. Xie, A. Vedaldi, and A. Zisserman, “VGGSound: A Large-Scale Audio-Visual Dataset,” inProc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), pp. 721–725, May 2020, doi: 10.1109/ICASSP40776.2020.9053174

  29. [30]

    Metabolic Cost as a Unifying Principle Governing Neuronal Biophysics,

    A. Hasenstaub, S. Otte, E. Callaway, and T. J. Sejnowski, “Metabolic Cost as a Unifying Principle Governing Neuronal Biophysics,”Proc. Natl. Acad. Sci. USA, vol. 107, no. 27, pp. 12329–12334, 2010, doi: 10.1073/pnas.0914886107. Available:https://www.pnas.org/doi/ abs/10.1073/pnas.0914886107

  30. [31]

    The Information Bottle- neck Method,

    N. Tishby, F. C. Pereira, and W. Bialek, “The Information Bottle- neck Method,” inProc. 37th Annu. Allerton Conf. Commun., Control Comput., pp. 368–377, 1999. Available:https://arxiv.org/abs/ physics/0004057

  31. [32]

    Neural Networks, Principal Components and Subspaces,

    E. Oja, “Neural Networks, Principal Components and Subspaces,” Int. J. Neural Syst., vol. 1, no. 1, pp. 61–68, 1989, doi: 10.1142/S0129065789000475

  32. [33]

    D. O. Hebb,The Organization of Behavior: A Neuropsychological The- ory. New York, NY, USA: John Wiley & Sons, 1949

  33. [34]

    Learning in Non-Linear Constrained Hebbian Networks,

    E. Oja, “Learning in Non-Linear Constrained Hebbian Networks,” in Proc. ICANN’91, pp. 385–390, 1991

  34. [35]

    Principal Component Analysis Learning Algorithms: A Neurobiological Anal- ysis,

    K. J. Friston, C. D. Frith, and R. S. J. Frackowiak, “Principal Component Analysis Learning Algorithms: A Neurobiological Anal- ysis,”Proc. Biol. Sci., vol. 254, no. 1339, pp. 47–54, 1993, doi: 10.1098/rspb.1993.0125

  35. [36]

    H. S. Lopez, B. Burger, R. Dickstein, N. L. Desmond, and W. B. Levy, “Associative Synaptic Potentiation and Depression: Quantification of Dissociable Modifications in the Hippocampal Dentate Gyrus Favors a Particular Class of Synaptic Modification Equations,”Synapse, vol. 5, no. 1, pp. 33–47, 1990, doi: 10.1002/syn.890050103

  36. [37]

    Synaptic Scaling—An Artificial Neu- ral Network Regularization Inspired by Nature,

    M. Hofmann and P. M¨ ader, “Synaptic Scaling—An Artificial Neu- ral Network Regularization Inspired by Nature,”IEEE Trans. Neu- ral Netw. Learn. Syst., vol. 33, no. 7, pp. 3094–3108, 2022, doi: 10.1109/TNNLS.2021.3050422

  37. [38]

    Auto-Encoding Variational Bayes,

    D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” in Proc. Int. Conf. Learn. Represent. (ICLR), 2014. Available:https: //arxiv.org/abs/1312.6114. 37

  38. [39]

    The Kinetics Human Action Video Dataset,

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman, “The Kinetics Human Action Video Dataset,” arXiv preprint arXiv:1705.06950, 2017, doi: 10.48550/arXiv.1705.06950

  39. [40]

    Audio-Visual Class- Incremental Learning,

    W. Pian, S. Mo, Y. Guo, and Y. Tian, “Audio-Visual Class- Incremental Learning,” inProc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), Oct. 2023, doi: 10.1109/ICCV51070.2023.00717

  40. [41]

    VideoMAE: Masked Au- toencoders Are Data-Efficient Learners for Self-Supervised Video Pre- Training,

    Z. Tong, Y. Song, J. Wang, and L. Wang, “VideoMAE: Masked Au- toencoders Are Data-Efficient Learners for Self-Supervised Video Pre- Training,” inAdv. Neural Inf. Process. Syst., vol. 35, 2022. Available: https://proceedings.neurips.cc/paper_files/paper/2022/ file/416f9cb3276121c42eebb86352a4354a-Paper-Conference. pdf

  41. [42]

    Masked Autoencoders That Listen,

    P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked Autoencoders That Listen,” inAdv. Neural Inf. Process. Syst., vol. 35, pp. 28708–28720, 2022. Available: https://proceedings.neurips.cc/paper_files/paper/2022/ file/b89d5e209990b19e33b418e14f323998-Paper.pdf

  42. [43]

    Performance-Optimized Hierarchical Models Predict Neural Responses in Higher Visual Cortex,

    D. L. K. Yamins, H. Hong, C. F. Cadieu, E. A. Solomon, D. Seibert, and J. J. DiCarlo, “Performance-Optimized Hierarchical Models Predict Neural Responses in Higher Visual Cortex,”Proc. Natl. Acad. Sci. USA, vol. 111, no. 23, pp. 8619–8624, 2014, doi: 10.1073/pnas.1403112111

  43. [44]

    A Task-Optimized Neural Network Replicates Human Auditory Behavior, Predicts Brain Responses and Reveals a Cortical Processing Hierarchy,

    A. J. E. Kell, D. L. K. Yamins, E. N. Shook, S. V. Norman-Haignere, and J. H. McDermott, “A Task-Optimized Neural Network Replicates Human Auditory Behavior, Predicts Brain Responses and Reveals a Cortical Processing Hierarchy,”Neuron, vol. 98, no. 3, pp. 630–644, 2018, doi: 10.1016/j.neuron.2018.03.044

  44. [45]

    Masked Autoencoders Are Scalable Vision Learners,

    K. He, X. Chen, S. Xie, Y. Li, P. Doll´ ar, and R. Girshick, “Masked Autoencoders Are Scalable Vision Learners,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 16000–16009, 2022. Available:https://openaccess.thecvf.com/content/CVPR2022/ html/He_Masked_Autoencoders_Are_Scalable_Vision_Learners_ CVPR_2022_paper.html

  45. [46]

    A Hierar- chy of Temporal Receptive Windows in Human Cortex,

    U. Hasson, E. Yang, I. Vallines, D. J. Heeger, and N. Rubin, “A Hierar- chy of Temporal Receptive Windows in Human Cortex,”J. Neurosci., vol. 28, no. 10, pp. 2539–2550, 2008, doi: 10.1523/JNEUROSCI.5487- 07.2008. 38

  46. [47]

    Slow Cortical Dynamics and the Accumulation of Information Over Long Timescales,

    C. J. Honey, T. Thesen, T. H. Donner, L. J. Silbert, C. E. Carlson, O. Devinsky, W. K. Doyle, N. Rubin, D. J. Heeger, and U. Hasson, “Slow Cortical Dynamics and the Accumulation of Information Over Long Timescales,”Neuron, vol. 76, no. 2, pp. 423–434, 2012, doi: 10.1016/j.neuron.2012.08.011

  47. [48]

    Normalization as a Canonical Neural Computation,

    M. Carandini and D. J. Heeger, “Normalization as a Canonical Neural Computation,”Nat. Rev. Neurosci., vol. 13, no. 1, pp. 51–62, 2012, doi: 10.1038/nrn3136

  48. [49]

    A Brain-Inspired Algorithm for Training Highly Sparse Neural Networks,

    Z. Atashgahi, J. Pieterse, S. Liu, D. C. Mocanu, R. Veldhuis, and M. Pechenizkiy, “A Brain-Inspired Algorithm for Training Highly Sparse Neural Networks,”Mach. Learn., vol. 111, no. 12, pp. 4411–4452, 2022, doi: 10.1007/s10994-022-06266-w

  49. [50]

    Learning Sparse Neu- ral Networks ThroughL 0 Regularization,

    C. Louizos, M. Welling, and D. P. Kingma, “Learning Sparse Neu- ral Networks ThroughL 0 Regularization,” inProc. Int. Conf. Learn. Represent. (ICLR), OpenReview.net, 2018. Available:https:// openreview.net/forum?id=H1Y8hhg0b

  50. [51]

    Pruning Neu- ral Networks Without Any Data by Iteratively Conserving Synaptic Flow,

    H. Tanaka, D. Kunin, D. Yamins, and S. Ganguli, “Pruning Neu- ral Networks Without Any Data by Iteratively Conserving Synaptic Flow,” inAdv. Neural Inf. Process. Syst., vol. 33, 2020

  51. [52]

    Sparse Networks from Scratch: Faster Training Without Losing Performance,

    T. Dettmers and L. Zettlemoyer, “Sparse Networks from Scratch: Faster Training Without Losing Performance,” arXiv preprint arXiv:1907.04840, 2019, doi: 10.48550/arXiv.1907.04840

  52. [53]

    Rigging the Lottery: Making All Tickets Winners,

    U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen, “Rigging the Lottery: Making All Tickets Winners,” inProc. Int. Conf. Mach. Learn. (ICML), vol. 119, pp. 2943–2952, Jul. 2020. Available:https: //proceedings.mlr.press/v119/evci20a.html

  53. [54]

    Braitenberg and A

    V. Braitenberg and A. Sch¨ uz,Cortex: Statistics and Geometry of Neu- ronal Connectivity. Berlin, Germany: Springer, 1998

  54. [55]

    Synaptic Pruning in Devel- opment: A Computational Account,

    G. Chechik, I. Meilijson, and E. Ruppin, “Synaptic Pruning in Devel- opment: A Computational Account,”Neural Comput., vol. 10, no. 7, pp. 1759–1777, 1998, doi: 10.1162/089976698300017124

  55. [56]

    Activation Learning by Local Competitions,

    H. Zhou, “Activation Learning by Local Competitions,” arXiv preprint arXiv:2209.13400, 2022, doi: 10.48550/arXiv.2209.13400

  56. [57]

    A Simplified Neuron Model as a Principal Component Analyzer,

    E. Oja, “A Simplified Neuron Model as a Principal Component Analyzer,”J. Math. Biol., vol. 15, no. 3, pp. 267–273, 1982, doi: 10.1007/BF00275687. 39

  57. [58]

    TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems,

    M. Abadiet al., “TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems,” arXiv preprint arXiv:1603.04467, 2016, doi: 10.48550/arXiv.1603.04467. Available:https://www.tensorflow. org/

  58. [59]

    Adam: A Method for Stochastic Optimiza- tion,

    D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimiza- tion,” inProc. Int. Conf. Learn. Represent. (ICLR), 2015. Available: https://hdl.handle.net/11245/1.505367

  59. [60]

    Individual Comparisons by Ranking Methods,

    F. Wilcoxon, “Individual Comparisons by Ranking Methods,”Biom. Bull., vol. 1, no. 6, pp. 80–83, 1945, doi: 10.2307/3001968

  60. [61]

    A Simple Sequentially Rejective Multiple Test Procedure,

    S. Holm, “A Simple Sequentially Rejective Multiple Test Procedure,” Scand. J. Stat., vol. 6, no. 2, pp. 65–70, 1979, doi: 10.2307/4615733

  61. [62]

    Is Bio-Inspired Learning Better Than Backprop? Benchmarking Bio Learning vs. Backprop,

    M. Gupta, S. K. Modi, H. Zhang, J. H. Lee, and J. H. Lim, “Is Bio-Inspired Learning Better Than Backprop? Benchmarking Bio Learning vs. Backprop,” arXiv preprint arXiv:2212.04614, 2023, doi: 10.48550/arXiv.2212.04614

  62. [63]

    Unsupervised Repre- sentation Learning with Hebbian Synaptic and Structural Plasticity in Brain-Like Feedforward Neural Networks,

    N. Ravichandran, A. Lansner, and P. Herman, “Unsupervised Repre- sentation Learning with Hebbian Synaptic and Structural Plasticity in Brain-Like Feedforward Neural Networks,”Neurocomputing, vol. 626, Art. no. 129440, 2025, doi: 10.1016/j.neucom.2025.129440

  63. [64]

    Supervised Hebbian learning in Deep Counterstream Associative Networks

    A. Knoblauch, “Supervised Hebbian Learning in Deep Counterstream Associative Networks,” arXiv preprint arXiv:2606.29528, 2026, doi: 10.48550/arXiv.2606.29528

  64. [65]

    Accelerating the Training of Feedforward Neural Networks Using Generalized Hebbian Rules for Initializing the Internal Representations,

    N. B. Karayiannis, “Accelerating the Training of Feedforward Neural Networks Using Generalized Hebbian Rules for Initializing the Internal Representations,”IEEE Trans. Neural Netw., vol. 7, no. 2, pp. 419– 426, 1996, doi: 10.1109/72.485677

  65. [66]

    Oja's plasticity rule overcomes several challenges of training neural networks under biological constraints

    N. Shervani-Tabar, M. A. Mirhoseini, and R. Rosenbaum, “Oja’s Plas- ticity Rule Overcomes Several Challenges of Training Neural Networks Under Biological Constraints,” arXiv preprint arXiv:2408.08408, 2025, doi: 10.48550/arXiv.2408.08408

  66. [67]

    Scalable Bio-Inspired Training of Deep Neural Networks with Fas- tHebb,

    G. Lagani, F. Falchi, C. Gennaro, H. Fassold, and G. Amato, “Scalable Bio-Inspired Training of Deep Neural Networks with Fas- tHebb,”Neurocomputing, vol. 595, p. 127867, May 2024, doi: 10.1016/j.neucom.2024.127867

  67. [68]

    Unsupervised Learning by Competing Hidden Units,

    D. Krotov and J. J. Hopfield, “Unsupervised Learning by Competing Hidden Units,”Proc. Natl. Acad. Sci. U.S.A., vol. 116, no. 16, pp. 7723–7731, 2019, doi: 10.1073/pnas.1820458116. 40

  68. [69]

    A Simple Frame- work for Contrastive Learning of Visual Representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Frame- work for Contrastive Learning of Visual Representations,” inProc. Int. Conf. Mach. Learn. (ICML), vol. 119, pp. 1597–1607, 2020. Available: https://proceedings.mlr.press/v119/chen20j.html

  69. [70]

    Momentum Contrast for Unsupervised Visual Representation Learning,

    K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum Contrast for Unsupervised Visual Representation Learning,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 9726–9735, Jun. 2020, doi: 10.1109/CVPR42600.2020.00975

  70. [71]

    MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining,

    X. Donget al., “MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining,” inProc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pp. 10995–11005, 2023. Available:https://openaccess.thecvf.com/content/CVPR2023/ papers/Dong_MaskCLIP_Masked_Self-Distillation_Advances_ Contrastive_Language-Image_Pretraining_CVPR_2023_paper. pdf

  71. [72]

    Self-Distilled Self-Supervised Representation Learning,

    J. Jang, S. Kim, K. Yoo, C. Kong, J. Kim, and N. Kwak, “Self-Distilled Self-Supervised Representation Learning,” inProc. IEEE/CVF Win- ter Conf. Appl. Comput. Vis. (WACV), pp. 2829–2839, 2023. Available:https://openaccess.thecvf.com/content/WACV2023/ html/Jang_Self-Distilled_Self-Supervised_Representation_ Learning_WACV_2023_paper.html

  72. [73]

    Synaptic Energy Use and Supply,

    J. J. Harris, R. Jolivet, and D. Attwell, “Synaptic Energy Use and Supply,”Neuron, vol. 75, no. 5, pp. 762–777, 2012, doi: 10.1016/j.neuron.2012.08.019

  73. [74]

    Difference Target Propagation,

    D.-H. Lee, S. Zhang, A. Fischer, and Y. Bengio, “Difference Target Propagation,” inProc. Joint Eur. Conf. Mach. Learn. Knowl. Discov. Databases, ser. Lecture Notes in Computer Science, pp. 498–515, 2015, doi: 10.1007/978-3-319-23528-8 31

  74. [75]

    Direct Feedback Alignment Provides Learning in Deep Neural Networks,

    A. Nøkland, “Direct Feedback Alignment Provides Learning in Deep Neural Networks,” inAdv. Neural Inf. Process. Syst., vol. 29, pp. 1037– 1045, 2016. Available:https://arxiv.org/abs/1609.01596

  75. [76]

    Size and Depth of Mono- tone Neural Networks: Interpolation and Approximation,

    D. Mikulincer and D. Reichman, “Size and Depth of Mono- tone Neural Networks: Interpolation and Approximation,” in Adv. Neural Inf. Process. Syst., vol. 35, 2022. Available:https: //proceedings.neurips.cc/paper_files/paper/2022/file/ 24c523085d10743633f9964e0623dbe0-Paper-Conference.pdf

  76. [77]

    Certified Monotonic Neural Networks,

    X. Liu, X. Han, N. Zhang, and Q. Liu, “Certified Monotonic Neural Networks,” inAdv. Neural Inf. Process. Syst., vol. 33, 2020. Avail- 41 able:https://proceedings.neurips.cc/paper_files/paper/ 2020/file/b139aeda1c2914e3b579aafd3ceeb1bd-Paper.pdf

  77. [78]

    Un- derstanding Encoder-Decoder Structures in Machine Learning Using Information Measures,

    J. F. Silva, V. Faraggi, C. Ram ´ ırez, A. Ega na, and E. Pavez, “Un- derstanding Encoder-Decoder Structures in Machine Learning Using Information Measures,”Signal Process., vol. 234, p. 109983, 2025, doi: 10.1016/j.sigpro.2025.109983

  78. [79]

    Information Dropout: Learning Optimal Representations Through Noisy Computation,

    A. Achille and S. Soatto, “Information Dropout: Learning Optimal Representations Through Noisy Computation,”IEEE Trans. Pat- tern Anal. Mach. Intell., vol. 40, no. 12, pp. 2897–2905, 2018, doi: 10.1109/TPAMI.2017.2784440

  79. [80]

    Learning Understandable Neu- ral Networks With Nonnegative Weight Constraints,

    J. Chorowski and J. M. Zurada, “Learning Understandable Neu- ral Networks With Nonnegative Weight Constraints,”IEEE Trans. Neural Netw. Learn. Syst., vol. 26, pp. 62–69, 2015, doi: 10.1109/TNNLS.2014.2310059

  80. [81]

    Optimal Unsupervised Learning in a Single-Layer Lin- ear Feedforward Neural Network,

    T. D. Sanger, “Optimal Unsupervised Learning in a Single-Layer Lin- ear Feedforward Neural Network,”Neural Networks, vol. 2, no. 6, pp. 459–473, 1989, doi: 10.1016/0893-6080(89)90044-0

Showing first 80 references.