Pith. sign in

Paper Citation Record · LEDGER

Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2208.03306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2208.03306 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T23:17:19.787815Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

26
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 93a51336-5131-4fde-a1f8-d295bbeafd7f · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:09:13.060337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:5546a687746a79a71fcdec5cdba792a20b5c8974494494e710c449df85cc6ee4

Observation 1a67f19a-4f70-4387-8c2c-84c82cd66125 · inbound

Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer cites this paper.

Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:13:30.125991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T22:11:59.206066Z digest=sha256:b11b8b1536813e5fae605780bea94a3223ef545ada85fb11d6d9e65f61c52a91

Observation 01bb3cee-ca4e-496c-a52b-66f6bff67e97 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.906658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:df7ce54efdfcdf81a9453a890efea03155c56b8e9ff0f9a42fb8f742ba0a2f51

Observation e1a6904e-6ec4-49ee-b3d0-09f681773735 · inbound

Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch cites this paper.

Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T23:17:19.787815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:17:19.787815Z digest=sha256:042a5a6acdeb62df8c9e694593cb2dc6ac6e3141e6d2817fc5da23a1dc762dd7

Observation 5bfd1ccc-4dbe-4e36-8663-d05a0b3951f0 · inbound

BTS: Harmonizing Specialized Experts into a Generalist LLM cites this paper.

BTS: Harmonizing Specialized Experts into a Generalist LLM Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T21:59:01.070134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:59:01.070134Z digest=sha256:676f763ba4ad7b40f44395bad38ab5ab0b23606ec268f4284c4d363d1d22001e

Observation 037a25a9-6e20-4cfb-b745-2e63c2eeec21 · inbound

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs cites this paper.

MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T17:01:41.122374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:01:41.122374Z digest=sha256:d2f585174e6014d8f749bd05d6117107deb51476b92380c4076a8540aa2d1ac7

Observation d91b4050-c3d9-4312-9c0b-14dd26beb3d9 · inbound

When One LLM Drools, Multi-LLM Collaboration Rules cites this paper.

When One LLM Drools, Multi-LLM Collaboration Rules Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T22:33:09.418737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:33:09.418737Z digest=sha256:d9f93eb8806148266e4995bd83413c563069fa588d72faa69edae5cd6ff821bc

Observation 204ed7b1-527f-4820-beff-4f0ec9fed988 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:00.952905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:32993479164c75d08499aa117971caf5ebd71297559a1fb27221772ef3bc512c

Observation 7b058177-784b-435a-a3ff-7acf76772da6 · inbound

Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging cites this paper.

Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.516966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:21.516966Z digest=sha256:685921c13efeaceb01619bab0450bd4d819efd676a04942cf47098d7290f557e

Observation b4b1e656-90ed-4aec-8020-cb5f269abe2c · inbound

NoLoCo: No-all-reduce Low Communication Training Method for Large Models cites this paper.

NoLoCo: No-all-reduce Low Communication Training Method for Large Models Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:02.441768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:02.441768Z digest=sha256:b37e00b94b8adcc4a0d60bffc42b036fc9ec40976ba4bf6c1d51eda3b7199614

Observation c4151b25-40eb-487f-81f6-58055bca8106 · inbound

Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model cites this paper.

Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:21.433548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:21.433548Z digest=sha256:eba052c3a6559349c313ddae6e76be6d9974f1516c3f087b7ffbb82dc9061cc8

Observation 3719c411-dbc3-4766-89de-cea2be10bd24 · inbound

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts cites this paper.

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:26.216747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:56:26.216747Z digest=sha256:949a7d90bcd0c9202dfe35cd590252cd49720df3c668f21f40679914b6e4e1eb

Observation 24e3a7d0-fe10-400a-8c73-3d34e3d9eb7e · inbound

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics cites this paper.

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:04.662169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:59:04.662169Z digest=sha256:f3fbf4c3db88f3184b04b70597dfa9716d4f860adb51190d18db168ba2ea6f16

Observation c829d2a0-d80c-476b-b417-ce674ea619e4 · inbound

FlexOlmo: Open Language Models for Flexible Data Use cites this paper.

FlexOlmo: Open Language Models for Flexible Data Use Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:16.051548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:16.051548Z digest=sha256:3325774b30dd3f2801466d2e00f52c585f97b7d3b6fc6d487bd2585d49f449c9

Observation 8f5852f7-d3e2-406d-bf99-b68c77247d1c · inbound

MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications cites this paper.

MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:48:15.103016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:47:51.509798Z digest=sha256:b6d26e96011702505e7ded744ad144df52b372bc5d1a28117dd67a76f3d4394d

Observation 27729a76-23bc-495a-aab5-fcad434f7204 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:56:11.347027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:39b0a2347abab0208cfa0d56686ef089443823de252e53c40ad6887383bf601a

Observation 4602cde6-ba2d-4709-b06d-03129ada670e · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 162

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:57:21.795138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:b912e53d8db14342f385040cb2bdbdb5539beeeb7573c7c16ae14aff218274dc

Observation 3dd25be9-893f-46ef-be1f-1dafa8d7d156 · inbound

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts cites this paper.

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:55:04.758253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:54:32.496951Z digest=sha256:de3c6b63860b9a42c2a79206dbc5b971dc498efd43628671110614c0fb500317

Observation 30bd7b97-324e-431d-9077-40337fc42078 · inbound

MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification cites this paper.

MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:23:31.881288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:21:10.641070Z digest=sha256:dbdac1fc117c35a1a55b07663c77cdad3e51c37a1cb370fa1bf99db8600ea3eb

Observation 6bef364a-0d99-4d45-9728-c51e1b4e1419 · inbound

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models cites this paper.

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:23:16.769512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:22:30.263086Z digest=sha256:70190d88bf654b5651b428b8c53ff667534552d4ad9379c7a583f68d8fd8fbae

Observation a6aac792-d8ae-41bb-b993-f2616357ffba · inbound

Decentralised AI Training and Inference with BlockTrain cites this paper.

Decentralised AI Training and Inference with BlockTrain Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.431608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T23:32:43.585170Z digest=sha256:b1aeeebd841c0a9e76db656bc796f6f100afd5ff02efbb2d611397bc6ffe5321

Observation cd94ee50-d4a0-4886-b0fe-3a2337f9ecdb · inbound

Decentralised AI Training and Inference with BlockTrain cites this paper.

Decentralised AI Training and Inference with BlockTrain Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T12:29:35.403453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:29:35.403453Z digest=sha256:0df6b3cd3e03fd1c18841cfbc896063d087653912911f64559625304779b6220

Observation 256c947c-611d-4a69-baa4-1ab1b55861fc · inbound

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield cites this paper.

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.043684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:46:40.510955Z digest=sha256:d445e74642bebd4fdd1110187dd573764182c0b9c9b4f9d9c0d349883c4d7d25

Observation 93570348-03da-426a-ba90-993117dbef4a · inbound

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield cites this paper.

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T09:24:14.388774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:24:14.388774Z digest=sha256:6f8a61e1f6686fb52dd10a5a02dd15b0bd89649e504a94b2f394e73c0f865dcc

Observation f25ca2e0-bf00-4380-9012-68d56b699d0d · inbound

Identifying Latent Concepts and Structures for Generalized Category Discovery cites this paper.

Identifying Latent Concepts and Structures for Generalized Category Discovery Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:03.350280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T14:42:01.334822Z digest=sha256:7b76529f09c430368dc9954bc2d9484ddb6d13dc5310a1a035e93e828a0dbedf

Observation f601a8d1-a5d2-4f76-965d-e75617e19f10 · inbound

Modular Foundation Models for Time-Series Perception in Digital Twins cites this paper.

Modular Foundation Models for Time-Series Perception in Digital Twins Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-12T01:22:51.284207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:22:51.284207Z digest=sha256:3fd8896cb24cb68aff2ccb40bb91113bbc46c7fa304ad70d56d44904e86c159c

Observation c2433688-1c4c-4442-a014-c6bea26815e8 · inbound

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization cites this paper.

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-08T00:41:22.247867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:41:22.247867Z digest=sha256:8394e96ef0e01e9e1688efa6c9aef99588aabf37e8093bae68a53eb45705c4ba