Pith. sign in

REVIEW 8 cited by

Modular Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.11529 v2 pith:GPNEEX6P submitted 2023-02-22 cs.LG

classification cs.LG
keywords learningmodelsmodularmodulestaskstransfercomputationdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transfer learning has recently become the dominant paradigm of machine learning. Pre-trained models fine-tuned for downstream tasks achieve better performance with fewer labelled examples. Nonetheless, it remains unclear how to develop models that specialise towards multiple tasks without incurring negative interference and that generalise systematically to non-identically distributed tasks. Modular deep learning has emerged as a promising solution to these challenges. In this framework, units of computation are often implemented as autonomous parameter-efficient modules. Information is conditionally routed to a subset of modules and subsequently aggregated. These properties enable positive transfer and systematic generalisation by separating computation from routing and updating modules locally. We offer a survey of modular architectures, providing a unified view over several threads of research that evolved independently in the scientific literature. Moreover, we explore various additional purposes of modularity, including scaling language models, causal inference, programme induction, and planning in reinforcement learning. Finally, we report various concrete applications where modularity has been successfully deployed such as cross-lingual and cross-modal knowledge transfer. Related talks and projects to this survey, are available at https://www.modulardeeplearning.com/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. Generalizable and Computational Efficient Channel Extrapolation for 6G: A Configurable AI-Driven Framework Built from a Modular Perspective

    eess.SP 2026-08 conditional novelty 6.0 of 10

    A three-stage modular AI framework, pretrain, cluster experts, and learn routing, improves channel extrapolation accuracy and cuts FLOPs in simulated 6G scenarios.

  2. DivMerge: A divergence-based model merging method for multi-tasking

    cs.LG 2025-09 conditional novelty 6.0 of 10

    DivMerge learns task-arithmetic merging weights by minimizing Jensen-Shannon divergence between each specialist model and the merged model, improving multi-task performance and scalability.

  3. Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Representation smoothness can be used to regularize training, stop early without validation labels, and guide active learning combined with parameter-efficient fine-tuning, reducing data and compute.

  4. Composing Linear Layers from Irreducibles

    cs.LG 2025-07 reject novelty 6.0 of 10

    A rotor-based layer built from bivector exponentials approximates LLM attention projections with O(log^2 d) parameters and competitive downstream performance.

  5. Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Adaptive systems can reduce costly core-parameter changes by first reorganizing access to existing computational modules — a principle called Accessibility Plasticity.

  6. Modular Foundation Models for Time-Series Perception in Digital Twins

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A gated bank of frozen self-supervised time-series encoders, aligned and aggregated by a Transformer, supports competitive multi-task perception for digital twins and hydro-generator virtual sensing.

  7. Compositional Learning for Modular Multi-Agent Self-Organizing Networks

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A two-tier agent framework with compositional reinforcement learning and predictive decision-making reduces handover failures, improves KPIs, and accelerates training in simulated self-organizing networks.

  8. DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer

    cs.CL 2025-05 conditional novelty 5.0 of 10

    DeFT-X applies SVD denoising to weight updates before magnitude pruning in composable sparse fine-tuning, showing small average gains over LT-SFT on NusaX and AmericasNLI.

Pith tools