Pith. sign in

REVIEW 4 cited by

Better together? Statistical learning in models made of modules

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1708.08719 v1 pith:KZ25MSGG submitted 2017-08-29 stat.ME

classification stat.ME
keywords modulesapproachesdatamodelmodelspropagationsettingsstatistical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In modern applications, statisticians are faced with integrating heterogeneous data modalities relevant for an inference, prediction, or decision problem. In such circumstances, it is convenient to use a graphical model to represent the statistical dependencies, via a set of connected "modules", each relating to a specific data modality, and drawing on specific domain expertise in their development. In principle, given data, the conventional statistical update then allows for coherent uncertainty quantification and information propagation through and across the modules. However, misspecification of any module can contaminate the estimate and update of others, often in unpredictable ways. In various settings, particularly when certain modules are trusted more than others, practitioners have preferred to avoid learning with the full model in favor of approaches that restrict the information propagation between modules, for example by restricting propagation to only particular directions along the edges of the graph. In this article, we investigate why these modular approaches might be preferable to the full model in misspecified settings. We propose principled criteria to choose between modular and full-model approaches. The question arises in many applied settings, including large stochastic dynamical systems, meta-analysis, epidemiological models, air pollution models, pharmacokinetics-pharmacodynamics, and causal inference with propensity scores.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Modelling Under-Reported Data: Pitfalls of Na\"ive Approaches and a New Statistical Framework for Epidemic Curve Reconstruction

    stat.AP 2025-09 conditional novelty 6.0 of 10

    A new framework approximates thinned count autoregressions with a normal-normal model and a latent Gaussian transform, enabling simple Bayesian reconstruction of under-reported epidemic curves.

  2. A Bayesian Spatio-Temporal Top-Down Framework for Estimating Opioid Use Disorder Risk Under Data Sparsity

    stat.AP 2025-06 conditional novelty 6.0 of 10

    A two-stage Bayesian downscaling model estimates county-level opioid use disorder risk from state-level survey counts, but its validation uses data generated by the model itself and shows large county-level errors.

  3. Bayesian Inference for Spatially-Temporally Misaligned Data Using Predictive Stacking

    stat.ME 2025-05 conditional novelty 6.0 of 10

    A modular Bayesian model with predictive stacking estimates how a spatially and temporally misaligned exposure relates to a block-level health outcome, demonstrated on California ozone and asthma emergency visits.

  4. Forecasting Influenza Hospitalizations Using a Bayesian Hierarchical Nonlinear Model with Discrepancy

    stat.AP 2024-12 conditional novelty 5.0 of 10

    A two-component Bayesian model, combining an asymmetric-Gaussian ILI forecast with a discrepancy term and a hospitalization regression, placed second among 21 non-ensemble models in the 2023-24 CDC FluSight challenge.

Pith tools