Pith. sign in

REVIEW 5 major objections 3 minor 64 references

Federated generative event models for tokenized electronic health records

T0 review · 5 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Federated training of tokenized generative event models nearly matches centralized training on pooled data, with most gains in 5-10 rounds.

desk verdict A real three-health-system evaluation of federated GEMs whose central claims largely hold; the MIMIC-frozen tokenizer is a legitimate caveat, not a fatal flaw. read the letter →

arxiv 2608.02939 v1 pith:IPUVL53X submitted 2026-08-03 cs.LG cs.CY

classification cs.LGcs.CY
keywords federatedlearninggenerativeeventmodelselectronichealthrecordsrepresentation-basedinferencecross-sitetransportabilityclinicalpredictionFedAvgtokenization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether hospitals can jointly train a generative model of patient event sequences without sharing records, and whether the resulting model is worth using at any single hospital. It reports that on 122,251 ICU stays from three independent health systems, a generative event model (GEM) trained to predict the next clinical event transfers across sites with a ROC-AUC penalty of 0.025, versus 0.079 for a gradient-boosted baseline, and that FedAvg and FedAvgM recover most of the performance of training on centrally pooled data within 5-10 communication rounds. The paper's main boundary condition is that multi-site data helps most when the target hospital has little local data: after a few thousand local examples, site-trained models catch up, and centralized pooling offers only modest gains over strong local models. A sympathetic reader would take the central claim to be that federated aggregation is technically solved for this setting and that the remaining bottleneck is learning representations that actually improve a specific target site.

What carries the argument

The load-bearing object is the tokenized generative event model (GEM): a 76.9-million-parameter transformer trained from scratch to predict the next token in a hospitalization's event sequence, where tokens encode demographics, transfers, lab orders and results, vitals, medications, and other care events in a shared ICU data format with numerical values binned into deciles. The companion mechanism is representation-based inference: the sequence is truncated at 24 hours, the model's final hidden layer is extracted as a fixed vector, and a logistic regression per outcome is trained on those vectors. Federated learning enters as weighted averaging of model weights (FedAvg, plus momentum and adaptive variants), with each site training on a one-nth fraction of its data per round; the measured saturation by 5-10 rounds is what makes the federated claim practical.

What would settle it

Re-tokenize each hospital's records with bins and vocabulary learned on that hospital's own training split, retrain GEMs, and recompute the cross-site transfer penalties and FedAvg-versus-centralized gaps; if the GEM transportability advantage shrinks or the federated gap changes materially, the paper's headline numbers are in part artifacts of the fixed tokenizer rather than properties of the model or federated algorithm.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that pretraining a transformer to predict the next token in a shared, expert-mapped clinical event sequence produces patient representations that are markedly more portable across hospitals than representations from conventional supervised models. This portability is quantified as a 0.025 ROC-AUC average transfer penalty for the extended-trained GEM, compared with 0.079 for gradient-boosted trees and 0.077 for logistic regression; the PR-AUC penalty is 0.027 versus 0.089. The paper also claims that federated averaging (FedAvg and FedAvgM) comes within 0.010-0.024 ROC-AUC of centralized GEM training, while FedAdam fails by 0.13-0.17 ROC-AUC, and that nearly all federated benefit appears within 5-10 rounds. It further claims that centralized multi-site training improves on complete local training by only small amounts (0.003-0.009 ROC-AUC for the optimized GEM), so the practical payoff of multi-site learning is concentrated at data-limited sites.

Load-bearing premise

The entire cross-site comparison assumes that a token vocabulary and numerical value bins learned on one hospital's training split and then frozen for the other two hospitals preserve clinical meaning equally at all sites, so that differences in performance reflect model transportability rather than tokenization mismatch.

Editorial extensions

If this is right

  • A hospital joining a federation with little local data can expect a usable head start from a model trained at other sites; after a few thousand local examples, models trained on local data alone become competitive.
  • Federated GEM training requires only about 5-10 communication rounds to capture most of the benefit, so communication overhead need not be a barrier.
  • Server-side adaptive optimization such as FedAdam is counterproductive for these models; simple weighted averaging with or without momentum is the right default.
  • The main remaining gap to centralized training is not the federated aggregation step but the limited transferability of heterogeneous multi-site data to a target institution.
  • Representations from generative pretraining, even with only one epoch of training, transfer across institutions far better than supervised baselines, though extended training raises absolute performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One testable consequence not pursued in the paper: if the cross-site penalty is driven by the frozen tokenizer, then learning site-specific value bins with a shared token semantics (or matching bins by rank across sites) should reduce the already-small GEM penalty further; this could be checked by retraining with per-site bins.
  • The 5-10 round saturation suggests communication schedules could be made adaptive, stopping federated updates once client weights stabilize, without losing accuracy; this is not tested in the paper.
  • Because only representation-based inference was used, the transportability conclusion may not extend to generative inference or supervised fine-tuning, where cross-site differences in documentation style could be amplified; that is an open question the paper acknowledges.
  • A practical deployment reading, implicit in the results, is that federated models should be positioned as onboarding tools for data-limited sites rather than permanent replacements for local training; if correct, procurement decisions should emphasize easy participation over massive federation scale.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 3 minor

Summary. The manuscript evaluates federated training of tokenized generative event models (GEMs) on intensive care data from three health systems (UCMC, NU, MIMIC) harmonized to CLIF-2.1, totaling 122,251 hospitalizations and 12 post-24-hour prediction tasks. Models are compared in within-site, cross-site, centralized, and federated (FedAvg, FedAvgM, FedAdam) configurations using representation-based inference. The principal claims are that GEM representations are substantially more transportable than LightGBM or logistic regression (cross-site ROC-AUC penalties of 0.025 versus 0.079), that FedAvg and FedAvgM recover most of the centralized multi-site performance (deficits of 0.010-0.024 ROC-AUC), that gains saturate within 5-10 communication rounds, and that multi-site training is most valuable when local data are limited.

Significance. If the quantitative claims hold, the paper provides a valuable multi-site benchmark separating federated optimization losses from representation transfer losses in a clinically realistic setting. The study has notable strengths: equal one-epoch training budgets for federated and centralized comparisons, per-outcome exclusion of patients with pre-24h events, bootstrap confidence intervals, a released code framework (coreopsis), and an evaluation design that does not fit constants to the held-out data. The main claims are falsifiable and of direct practical interest for ICU prediction. However, the magnitude of several headline numbers rests on a MIMIC-frozen tokenizer and on single-seed comparisons of small differences, so the quantitative conclusions need additional robustness work.

major comments (5)
  1. [Section 3.2, Figure B1] The tokenizer and decile bins are learned exclusively from MIMIC training data and then frozen for UCMC and NU (Section 3.2). Because every cross-site and federated comparison runs through this shared token space, the transfer penalties in Table 4 and the federated deficits in Table 5 could partly reflect token-semantic misalignment between MIMIC-derived bins and UCMC/NU raw values rather than model transportability. Figure B1 documents substantial cross-site differences in quantile-token distributions, and the limitations paragraph concedes that fixing MIMIC bins 'may have affected institutions differently.' I request a robustness analysis that re-learns bins per site or on a pooled sample and reports whether the GEM transportability advantage and the FedAvg deficits persist.
  2. [Section 4.2, Tables 3-4] The claim that centralized multi-site training yields only modest improvements is stated without significance testing. The site-specific ROC-AUC gains for GEM-* are 0.004 at UCMC, 0.009 at NU, and 0.003 at MIMIC, all well within the reported bootstrap confidence intervals of the local models. Please add paired bootstrap tests (as is already done for FedAvg versus FedAvgM) for centralized versus local training, at least for the headline ROC-AUC comparisons, and report the resulting p-values.
  3. [Section 4.4] The crossover claim that local models become competitive after 'approximately 3,000 training examples' is not accompanied by an estimator, confidence interval, or definition of competitiveness. Please define a margin (e.g., within 0.005 ROC-AUC of the multi-site comparator), estimate the crossover from the learning curves, and provide uncertainty for it.
  4. [Table A6 (hyponatremia)] The FedAdam PR-AUC values for hyponatremia (0.501, 0.501, 0.422) are one to two orders of magnitude larger than those of every other model for the same outcome (typically 0.01-0.05). This appears to be an error in the appendix table or in the FedAdam evaluation, and it will distort the aggregate FedAdam PR-AUC deficits reported in Table 5. Please correct the anomaly and re-run the FedAdam comparisons.
  5. [Sections 3.3-3.5] All results are based on a single training run per configuration, so the confidence intervals reflect only patient-level resampling and not training stochasticity. For a 76.9M-parameter transformer, the key differences (FedAvg/FedAvgM deficits of 0.010-0.024 ROC-AUC, centralized gains of 0.003-0.009) are small enough that seed variation may change the qualitative conclusions. Please report at least three seeds for the central configurations, or demonstrate that seed-to-seed variability is negligible.
minor comments (3)
  1. [Section 5 (limitations)] The limitations paragraph contains the sentence 'The authors are incredibly grateful to the teams at Beth-Israel/MIT, Northwestern Medicine, and UCMC who have made data centralization possible for this experiment,' which reads as an acknowledgment and should be moved to the Acknowledgments section.
  2. [Figure B1 caption] The caption mentions 'mean Wasserstein-1 distance between sites' without defining how the distance is computed over quantile-token distributions; please specify the exact comparison procedure.
  3. [Table 5 and Section 4.3] The statement that no significant difference between FedAvg and FedAvgM was detected (p>0.25 for all six comparisons) is not accompanied by details of the bootstrap test or the p-values; please report the procedure and the full set of p-values.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the central claims are direct empirical evaluations on held-out data, and the MIMIC-learned tokenizer is an acknowledged preprocessing choice rather than a fitted predictor from which the results are derived.

full rationale

The paper's main claims—GEM transfer penalties of 0.025 ROC-AUC vs. 0.079 for LightGBM, FedAvg/FedAvgM deficits of 0.010–0.024 ROC-AUC relative to centralized GEM training, and performance saturation after 5–10 rounds—are computed directly from held-out institutional splits (Section 3.5; Tables 3–6). The only fitted quantities are standard downstream probe hyperparameters (LR L2 strength via 5-fold CV) and the tokenizer's decile bins learned on MIMIC training data. Section 3.2 states: 'We then froze the learned vocabulary and bins and applied the tokenizer to the tuning and held-out splits in MIMIC, and all splits in UCMC and NU.' That frozen tokenizer is a preprocessing constant applied identically to all models; it is not the quantity being predicted, and LightGBM and LR consume the same tokenized features, so the relative GEM-versus-baseline comparisons are not forced by construction. The paper explicitly acknowledges the tokenizer limitation: 'The vocabulary and numerical bins were learned from MIMIC and then fixed across sites, which enabled a common token space but may have affected institutions differently.' This is a robustness or confounding concern, not circularity: no equation reduces a claimed result to a fitted parameter, and no load-bearing premise is justified only by the authors' own prior work. The self-citations (e.g., refs. 6, 16, 21, 28) describe methods, prior GEM workflows, and benchmarks; they are not invoked as uniqueness theorems or as substitutes for the empirical transfer and federated results. Therefore the derivation is self-contained against external held-out benchmarks, and the circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central empirical claims rest on a modest set of design choices: the MIMIC-derived tokenizer, fixed epoch and round budgets, and standard assumptions about CLIF harmonization and temporal split exchangeability. No new theoretical entities or fitted physical constants are introduced. The chosen hyperparameters affect absolute performance but are held constant across the model families being compared, so they do not by themselves create the transportability advantage.

free parameters (3)
  • Decile cutoffs for category-value tokens = Not reported (defines the 1344-token vocabulary)
    These cutoffs are fit to the MIMIC training split and frozen across all sites. The cross-site comparability claims depend on this quantization being appropriate for UCMC and NU.
  • Number of federated communication rounds = 10
    Chosen for the headline federated comparison in Section 4.3. The paper's own saturation analysis shows most gains by 5-10 rounds, so this choice does not inflate the federated result.
  • Training epochs = 1 (GEM, federated, centralized one-epoch); 5 (GEM-*)
    The federated-versus-centralized comparison uses a one-epoch budget for all GEMs. The headline transportability numbers use the five-epoch GEM-*. Results may depend on these budgets.
assumptions (4)
  • domain assumption CLIF-2.1 mappings faithfully align site-specific ICU codes into a shared semantic space.
    Cross-site comparisons assume tokens have comparable meaning across institutions; lossy mappings would mix semantic mismatch with model transportability. Invoked in Section 3.2.
  • domain assumption MIMIC-learned vocabulary and decile bins are adequate for UCMC and NU data.
    Unseen events or skewed bins could distort non-MIMIC representations; Figure B1 shows substantial cross-site bin differences, and the limitations section acknowledges the risk.
  • domain assumption Temporal splits for UCMC and NU are exchangeable with random splits.
    UCMC and NU are split by admission time while MIMIC is random; time-varying practice or population shifts could bias held-out estimates (Section 3.1).
  • domain assumption Next-token pretraining plus logistic regression probing is a valid way to measure representation quality.
    All GEM conclusions use representation-based inference (Section 3.5); they are not established for generative or supervised fine-tuning.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Federated generative event models for tokenized electronic health records." pith.science (2026). https://pith.science/paper/IPUVL53X

@misc{pith2026260802939,
  author       = {Pith},
  title        = {Pith review of: Federated generative event models for tokenized electronic health records},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IPUVL53X}},
  note         = {Machine review of arXiv:2608.02939}
}
read the original abstract

Electronic health record foundation models are limited by institutionally siloed data and substantial performance degradation under cross-site transfer. We evaluated federated training of tokenized generative event models (GEMs) across 122,251 intensive care hospitalizations from three independent health systems harmonized to the Common Longitudinal ICU Data Format. Models were assessed on 12 post-24-hour clinical prediction tasks using within-site, cross-site, centralized, and federated training configurations. GEMs achieved the highest mean within-site and cross-site ROC-AUC and were substantially more transportable than conventional supervised models: their average cross-site penalties were 0.025 ROC-AUC and 0.027 PR-AUC, compared with 0.079 and 0.089 for LightGBM. Federated Learning (FedAvg and FedAvgM) approached the performance of centralized GEM training, with most gains obtained within 5-10 communication rounds. However, centralized multi-site training provided only modest improvements over complete local training. Multi-site models were most useful when local training data were limited, with their advantage narrowing as institutional data accumulated. These findings show that federated GEM training is technically feasible and preserves most centralized performance, but that the main open challenge is learning transportable representations to translate larger, but heterogeneous data from multiple health systems into a reliable target-site benefit.

Figures

Figures reproduced from arXiv: 2608.02939 by the authors.

Figure 1
Figure 1. Experimental design: EHR data is converted to the CLIF-2.1 standard in parallel for each of the three sites. A cocoa-tokenizer instance is learned on MIMIC training data and then applied to the other datasets to convert patient records into sequences of tokens. Training and tuning sequences are used at each respective site by a cotorra instance to train a GEM from scratch. The coreopsis software introduced in this p… view at source ↗
Figure 2
Figure 2. ROC-AUC performance on each respective fixed test set vs. number of training examples [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 53 canonical work pages

  1. [1]

    Scaling laws for neural language models

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models.” arXiv:2001.08361, 2020

  2. [2]

    An empirical analysis of compute-optimal large language model training,

    J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre, “An empirical analysis of compute-optimal large language model training,” inAdv....

  3. [3]

    Exploring scaling laws for EHR foundation models

    S. Zhang, Q. Liu, N. Usuyama, C. Wong, T. Naumann, and H. Poon, “Exploring scaling laws for EHR foundation models.” arXiv:2505.22964, 2025

  4. [4]

    Generative medical event models improve with scale

    S. Waxler, P. Blazek, D. White, D. Sneider, K. Chung, M. Nagarathnam, P. Williams, H. Voeller, K. Wong, M. Swanhorst, S. Zhang, N. Usuyama, C. Wong, T. Naumann, H. Poon, A. Loza, D. Meeker, S. Hain, and R. Shah, “Generative medical event models improve with scale.” arXiv:2508.12104, 2025

  5. [5]

    A multi-center study on the adaptability of a shared foundation model for electronic health records,

    L. L. Guo, J. Fries, E. Steinberg, S. L. Fleming, K. Morse, C. Aftandilian, J. Posada, N. Shah, and L. Sung, “A multi-center study on the adaptability of a shared foundation model for electronic health records,”npj Digit. Med., vol. 7, 2024

  6. [6]

    Foundation models for electronic health records: representation dynamics and transferability

    M. C. Burkhart, B. Ramadan, Z. Liao, K. Chhikara, J. C. Rojas, W. F. Parker, and B. K. Beaulieu-Jones, “Foundation models for electronic health records: representation dynamics and transferability.” arXiv:2504.10422, 2025

  7. [7]

    FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records

    C. Pang, V. Jeanselme, Y. S. Choi, X. Jiang, Z. Jing, A. Kashyap, Y. Kobayashi, Y. Li, F. Pollet, K. Natarajan, and S. Joshi, “FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records.” arXiv:2505.16941, 2025

  8. [8]

    Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2),

    S. N. Murphy, G. Weber, M. Mendis, V. Gainer, H. C. Chueh, S. Churchill, and I. Kohane, “Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2),” J. Am. Med. Inform. Assoc., vol. 17, no. 2, 2010

Show all 64 references
  1. [9]

    Feasibility and utility of applications of the common data model to multiple, disparate observational health databases,

    E. A. Voss, R. Makadia, A. Matcho, Q. Ma, C. Knoll, M. Schuemie, F. J. DeFalco, A. Londhe, V. Zhu, and P. B. Ryan, “Feasibility and utility of applications of the common data model to multiple, disparate observational health databases,”J. Am. Med. Inform. Assoc., vol. 22, no. 3, 2015

  2. [10]

    Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data,

    C. Hong, E. Rush, M. Liu, D. Zhou, J. Sun, A. Sonabend, V. M. Castro, P. Schubert, V. A. Panickan, T. Cai, L. Costa, Z. He, N. Link, R. Hauser, J. M. Gaziano, S. N. Murphy, G. Ostrouchov, Y.-L. Ho, E. Begoli, J. Lu, K. Cho, K. P. Liao, and T. Cai, “Clinical knowledge extractio...

  3. [11]

    Representation learning to advance multi-institutional studies with electronic health record data from US and France,

    D. Zhou, H. Tong, L. Wang, S. Liu, X. Xiong, Z. Gan, G. Romain, B. P. Hejblum, Y.-C. Liu, C. Hong, C.-L. Bonzel, T. Cai, K. Pan, Y.-L. Ho, L. Costa, V. A. Panickan, J. M. Gaziano, K. D. Mandl, V. Jouhet, R. Thiebaut, Z. Xia, K. Cho, K. Liao, and T. Cai, “Representation learnin...

  4. [12]

    Trends in ransomware attacks on US hospitals, clinics, and other health care delivery organizations, 2016-2021,

    H. T. Neprash, C. C. McGlave, D. A. Cross, B. A. Virnig, M. A. Puskarich, J. D. Huling, A. Z. Rozenshtein, and S. S. Nikpay, “Trends in ransomware attacks on US hospitals, clinics, and other health care delivery organizations, 2016-2021,”JAMA Health Forum, vol. 3, no. 12, 2022

  5. [13]

    Ransomware attacks and data breaches in US health care systems,

    J. X. Jiang, J. S. Ross, and G. Bai, “Ransomware attacks and data breaches in US health care systems,”JAMA Netw. Open, vol. 8, no. 5, 2025

  6. [14]

    A common longitudinal intensive care unit data format (CLIF) for critical illness research,

    J. C. Rojas, P. G. Lyons, K. Chhikara, V. Chaudhari, S. V. Bhavani, M. Nour, K. G. Buell, K. D. Smith, C. A. Gao, S. Amagai,et al., “A common longitudinal intensive care unit data format (CLIF) for critical illness research,”Intensive Care Med., vol. 51, 2025

  7. [15]

    Federation, not centralization: a new paradigm for electronic health record–based critical care research,

    P. G. Lyons, K. G. Buell, K. A. Connell, M. A. Christensen, C. H. Hochberg, S. Jain, W. F. Parker, K. Chhikara, J. C. Rojas, C. Blebea, S. V. Bhavani, A. K. Barker, N. Mesfin, N. E. Ingraham, and C. A. Gao, “Federation, not centralization: a new paradigm for electronic health ...

  8. [16]

    Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models,

    M. C. Burkhart, B. Ramadan, L. Solo, W. F. Parker, and B. K. Beaulieu-Jones, “Quantifying surprise in clinical care: Detecting highly informative events in electronic health records with foundation models,” inPac. Symp. Biocomput., vol. 31, World Scientific, 2026

  9. [17]

    Everyquery: Zero-shot clinical prediction via task-conditioned pretraining over electronic health records

    P. Chandak, G. Kondas, L. A. Friedman, I. Kohane, and M. McDermott, “Everyquery: Zero-shot clinical prediction via task-conditioned pretraining over electronic health records.” arXiv:2603.07900, 2026

  10. [18]

    Systematic review of foundation models for structured electronic health records,

    L. L. Guo, S. E. Arciniegas, A. P. Yan, J. Fries, G. A. Tomlinson, and L. Sung, “Systematic review of foundation models for structured electronic health records,”J. Am. Med. Inform. Assoc., vol. 33, no. 6, 2026

  11. [19]

    EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models,

    M. Wornow, R. Thapa, E. Steinberg, J. A. Fries, and N. Shah, “EHRSHOT: An EHR benchmark for few-shot evaluation of foundation models,” inNeurips Datasets and Benchmarks Track, vol. 36, 2023

  12. [20]

    Context clues: Evaluating long context models for clinical prediction tasks on EHRs,

    M. Wornow, S. Bedi, M. A. F. Hernandez, E. Steinberg, J. Fries, C. R´ e, O. Koyejo, and N. H. Shah, “Context clues: Evaluating long context models for clinical prediction tasks on EHRs,” in ICLR, 2025

  13. [21]

    Rep- resentation before training: A fixed-budget benchmark for generative medical event models

    I. Lee, L. Solo, M. C. Burkhart, B. Ramadan, W. F. Parker, and B. K. Beaulieu-Jones, “Rep- resentation before training: A fixed-budget benchmark for generative medical event models.” arXiv:2604.16775, 2026

  14. [22]

    Tokenization tradeoffs in structured EHR foundation models

    L. L. Guo, S. E. Arciniegas, J. J. Lee, A. P. Yan, G. Tomlinson, J. Fries, and L. Sung, “Tokenization tradeoffs in structured EHR foundation models.” arXiv:2603.15644, 2026

  15. [23]

    A multimodal and temporal foundation model for virtual patient representations at healthcare system scale

    A. Zhang, T. Ding, S. J. Wagner, C. Tian, M. Y. Lu, R. Pettit, J. E. Lewis, A. Misrahi, D. Mo, L. P. Le, and F. Mahmood, “A multimodal and temporal foundation model for virtual patient representations at healthcare system scale.” arXiv:2604.18570, 2026

  16. [24]

    Event stream GPT: A data pre-processing and modeling library for generative, pre-trained transformers over continuous-time sequences of complex events,

    M. McDermott, B. Nestor, P. Argaw, and I. S. Kohane, “Event stream GPT: A data pre-processing and modeling library for generative, pre-trained transformers over continuous-time sequences of complex events,” inAdv. Neur. Inf. Process. Syst., vol. 36, 2023

  17. [25]

    Foresight-a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study,

    Z. Kraljevic, D. Bean, A. Shek, R. Bendayan, H. Hemingway, J. A. Yeung, A. Deng, A. Baston, J. Ross, E. Idowu, J. T. Teo, and R. J. B. Dobson, “Foresight-a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective mod...

  18. [26]

    Zero shot health trajectory prediction using transformer,

    P. Renc, Y. Jia, A. E. Samir, J. Was, Q. Li, D. W. Bates, and A. Sitek, “Zero shot health trajectory prediction using transformer,”npj Digit. Med., vol. 7, 2024

  19. [27]

    Foundation model of electronic medical records for adaptive risk estimation,

    P. Renc, M. K. Grzeszczyk, N. Oufattole, D. Goode, Y. Jia, S. Bieganski, M. B. A. McDermott, J. Was, A. E. Samir, J. W. Cunningham, D. W. Bates, and A. Sitek, “Foundation model of electronic medical records for adaptive risk estimation,”GigaScience, vol. 14, 2025

  20. [28]

    Efficient generative prediction for EHR foundation models: The SCOPE and REACH estimators

    L. Solo, M. B. A. McDermott, W. F. Parker, B. Ramadan, M. C. Burkhart, and B. K. Beaulieu- Jones, “Efficient generative prediction for EHR foundation models: The SCOPE and REACH estimators.” arXiv:2602.03730, 2026

  21. [29]

    MOTOR: A time-to-event foundation model for structured medical records,

    E. Steinberg, J. A. Fries, Y. Xu, and N. Shah, “MOTOR: A time-to-event foundation model for structured medical records,” inICLR, 2024

  22. [30]

    EHRMamba: Towards generalizable and scalable foundation models for electronic health records,

    A. Fallahpour, M. Alinoori, W. Ye, X. Cao, A. Afkanpour, and A. Krishnan, “EHRMamba: Towards generalizable and scalable foundation models for electronic health records,” inML4H, vol. PMLR 259, 2025

  23. [31]

    Federated machine learning in healthcare: A systematic review on clinical applications and technical architecture,

    Z. L. Teo, L. Jin, N. Liu, S. Li, D. Miao, X. Zhang, W. Y. Ng, T. F. Tan, D. M. Lee, K. J. Chua, J. Heng, Y. Liu, R. S. Mong Goh, and D. S. Wei Ting, “Federated machine learning in healthcare: A systematic review on clinical applications and technical architecture,”Cell Rep. M...

  24. [32]

    Federated learning authenticity standard for healthcare as derived from lessons in self-driving cars,

    R. Santos and P. A. Keane, “Federated learning authenticity standard for healthcare as derived from lessons in self-driving cars,”Commun. Med., 2026

  25. [33]

    MIMIC-IV, a freely accessible electronic health record dataset,

    A. E. W. Johnson, L. Bulgarelli, L. Shen, A. Gayles, A. Shammout, S. Horng, T. J. Pollard, S. Hao, B. Moody, B. Gow, L.-W. H. Lehman, L. A. Celi, and R. G. Mark, “MIMIC-IV, a freely accessible electronic health record dataset,”Sci. Data, vol. 10, 2023

  26. [34]

    The eICU collaborative research database, a freely available multi-center database for critical care research,

    T. J. Pollard, A. E. W. Johnson, J. D. Raffa, L. A. Celi, R. G. Mark, and O. Badawi, “The eICU collaborative research database, a freely available multi-center database for critical care research,” Sci. Data, vol. 5, no. 1, 2018

  27. [35]

    Federated learning for electronic health records,

    T. K. Dang, X. Lan, J. Weng, and M. Feng, “Federated learning for electronic health records,” ACM Trans. Intell. Syst. Technol., vol. 13, no. 5, 2022

  28. [36]

    Communication- efficient learning of deep networks from decentralized data,

    B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. Ag¨ uera y Arcas, “Communication- efficient learning of deep networks from decentralized data,” inAISTATS, vol. 54, 2017

  29. [37]

    Measuring the effects of non-identical data distribution for federated visual classification,

    H. Hsu, H. Qi, and M. Brown, “Measuring the effects of non-identical data distribution for federated visual classification,” inNeurips Workshop on Federated Learning, 2019

  30. [38]

    Federated learning of medical concepts embedding using BEHRT,

    O. Ben Shoham and N. Rappoport, “Federated learning of medical concepts embedding using BEHRT,”JAMIA Open, vol. 7, no. 4, 2024

  31. [39]

    Federated timeline syn- thesis: Scalable and private methodology for model training and deployment

    P. Renc, M. K. Grzeszczyk, L. Qian, N. Oufattole, J. Rasley, and A. Sitek, “Federated timeline syn- thesis: Scalable and private methodology for model training and deployment.” arXiv:2506.23358, 2025

  32. [40]

    Validation of a common data model for active safety surveillance research,

    J. M. Overhage, P. B. Ryan, C. G. Reich, A. G. Hartzema, and P. E. Stang, “Validation of a common data model for active safety surveillance research,”J. Am. Med. Inform. Assoc., vol. 19, no. 1, 2012

  33. [41]

    Federated learning for heterogeneous electronic health record systems with cost effective participant selection

    J. Kim, J. Kim, K. Hur, and E. Choi, “Federated learning for heterogeneous electronic health record systems with cost effective participant selection.” arXiv:2404.13318, 2026

  34. [42]

    PORTER: Language-grounded event representa- tions for portable structured ehr foundation models

    L. L. Guo, A. P. Yan, E. Vettese, and L. Sung, “PORTER: Language-grounded event representa- tions for portable structured ehr foundation models.” arXiv:2606.24102, 2026

  35. [43]

    Representation learning of structured data for medical foundation models,

    V. P. Dwivedi, V. Schlegel, A. T. Liu, T.-T. Nguyen, A. R. Kashyap, J. Wei, W.-H. Yin, S. Winkler, and R. T. Tan, “Representation learning of structured data for medical foundation models,” in UniReps, vol. PMLR 285, 2024

  36. [44]

    Continuous kidney replacement therapies: Core curriculum 2025,

    J. P. Teixeira, S. Hiremath, A. O. Kabli, O. G. Rewa, and E. G. Clark, “Continuous kidney replacement therapies: Core curriculum 2025,”Am. J. Kidney Dis., vol. 85, no. 6, 2025

  37. [45]

    Hohmann, F

    F. Hohmann, F. Fichtner, T. Becher, D. Schaedler, C. Putensen, T. Muders, I. Schroeder, C. Karagiannidis, H. Wrigge, D. Berger, M. Grupp, F. Grundeis, V. Buenger, A. Sachkova, S. Henkel, M. Habicher, M. Sander, S. Laudi, S. Weber-Carstens, and O. Moerer, “Clinical guideline fo...

  38. [46]

    Association between do not resuscitate/do not intubate status and resident physician decision-making: A national survey,

    E. K. Stevenson, H. M. Mehter, A. J. Walkey, and R. S. Wiener, “Association between do not resuscitate/do not intubate status and resident physician decision-making: A national survey,” Ann. Am. Thorac. Soc., vol. 14, no. 4, 2017

  39. [47]

    Prone position in ARDS patients: why, when, how and for whom,

    C. Gu´ erin, R. K. Albert, J. Beitler, L. Gattinoni, S. Jaber, J. J. Marini, L. Munshi, L. Papazian, A. Pesenti, A. Vieillard-Baron, and J. Mancebo, “Prone position in ARDS patients: why, when, how and for whom,”Intensive Care Med., vol. 46, no. 12, 2020

  40. [48]

    The third international consensus definitions for sepsis and septic shock (Sepsis-3),

    M. Singer, C. S. Deutschman, C. W. Seymour, M. Shankar-Hari, D. Annane, M. Bauer, R. Bellomo, G. R. Bernard, J.-D. Chiche, C. M. Coopersmith, R. S. Hotchkiss, M. M. Levy, J. C. Marshall, G. S. Martin, S. M. Opal, G. D. Rubenfeld, T. van der Poll, J.-L. Vincent, and D. C. Angus...

  41. [49]

    CEHR- BERT: Incorporating temporal information from structured EHR data to improve prediction tasks,

    C. Pang, X. Jiang, K. S. Kalluri, M. Spotnitz, R. Chen, A. Perotte, and K. Natarajan, “CEHR- BERT: Incorporating temporal information from structured EHR data to improve prediction tasks,” inML4H, 2021

  42. [50]

    Rethinking tokenization for clinical time series: When less is more,

    R. A. Attrach, R. Fani, D. Restrepo, Y. Jia, and P. Sch¨ uffler, “Rethinking tokenization for clinical time series: When less is more,” inML4H, 2025

  43. [51]

    The Llama 3 herd of models

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle,et al., “The Llama 3 herd of models.” arXiv 2407.21783, 2024

  44. [52]

    Decoupled weight decay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” inICLR, 2019

  45. [53]

    Adam: A method for stochastic optimization,

    D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” inICLR, 2015

  46. [54]

    Comparing biases for minimal network construction with back- propagation,

    S. Hanson and L. Pratt, “Comparing biases for minimal network construction with back- propagation,” inAdv. Neur. Inf. Process. Syst., 1988

  47. [55]

    RoFormer: Enhanced transformer with rotary position embedding,

    J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “RoFormer: Enhanced transformer with rotary position embedding,”Neurocomput., vol. 568, 2024

  48. [56]

    NEFTune: Noisy embeddings improve instruction finetuning,

    N. Jain, P. yeh Chiang, Y. Wen, J. Kirchenbauer, H.-M. Chu, G. Somepalli, B. R. Bartoldson, B. Kailkhura, A. Schwarzschild, A. Saha, M. Goldblum, J. Geiping, and T. Goldstein, “NEFTune: Noisy embeddings improve instruction finetuning,” inICLR, 2024

  49. [57]

    A method of solving a convex programming problem with convergence rate o(1/sqr(k)),

    Y. E. Nesterov, “A method of solving a convex programming problem with convergence rate o(1/sqr(k)),”Sov. Math. Dokl., vol. 27, 1983

  50. [58]

    Flower: A friendly federated learning research framework

    D. J. Beutel, T. Topal, A. Mathur, X. Qiu, J. Fernandez-Marques, Y. Gao, L. Sani, K. H. Li, T. Parcollet, P. P. B. de Gusm˜ ao, and N. D. Lane, “Flower: A friendly federated learning research framework.” arXiv:2007.14390, 2020

  51. [59]

    Some methods of speeding up the convergence of iteration methods,

    B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,”USSR Comput. Math. Math. Phys., vol. 4, no. 5, 1964

  52. [60]

    Adaptive federated optimization,

    S. Reddi, Z. Charles, M. Zaheer, Z. Garrett, K. Rush, J. Koneˇ cn´ y, S. Kumar, and H. B. McMahan, “Adaptive federated optimization,” inICLR, 2021

  53. [61]

    A closer look at AUROC and AUPRC under class imbalance,

    M. B. McDermott, H. Zhang, L. H. Hansen, G. Angelotti, and J. Gallifant, “A closer look at AUROC and AUPRC under class imbalance,” inAdv. Neur. Inf. Proc. Sys., 2024

  54. [62]

    Efron and R

    B. Efron and R. J. Tibshirani,An Introduction to the Bootstrap, vol. 57 ofMonographs on Statistics and Applied Probability. Chapman and Hall, 1993

  55. [63]

    Lightgbm: A highly efficient gradient boosting decision tree,

    G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, “Lightgbm: A highly efficient gradient boosting decision tree,” inAdv. Neur. Inf. Proc. Sys., vol. 30, Curran Associates, Inc., 2017

  56. [64]

    Physiobank, physiotoolkit, and physionet,

    A. L. Goldberger, L. A. N. Amaral, L. Glass, J. M. Hausdorff, P. C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley, “Physiobank, physiotoolkit, and physionet,” Circulation, vol. 101, no. 23, 2000. Appendix A. Supplementary tables Table A1.Breakdow...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.