Pith. sign in

Paper Citation Record · LEDGER

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging

As of 10 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2502.01804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01804 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:57.561399Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8853291-31a6-4438-838c-e8177ea277fc · outbound

This paper cites Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.391563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.391563Z digest=sha256:639ad6df2b4c441c97716400eecbe7a1a35ab7ad32866df58b814c2ad20264d1

Observation 36c7597e-646b-4b7d-9bcc-b4aa6344e198 · outbound

This paper cites Ensemble of averages: Improving model selection and boosting performance in domain generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Ensemble of averages: Improving model selection and boosting performance in domain generalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.396007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.396007Z digest=sha256:dd8ee248206b2b43a002817dd3e938a9092ef37c2e4ee15d616f03a2dfdacf0e

Observation dd5c57d8-ee63-42bb-95dd-e47857dbbeb0 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.400087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.400087Z digest=sha256:892f2fcd54808b1195c072d02c0e84c7be6282b9260ba2ebec1bd8941ba9f458

Observation f408d395-f482-4108-a0ad-dba645e9922d · outbound

This paper cites an unresolved cited work.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.404425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.404425Z digest=sha256:154b300a9c41a558b828ff06516ed9a32d920cd90e1741b98f1a4af9bdb0951b

Observation feb98c8f-41b0-4247-8458-6caf21c33376 · outbound

This paper cites Fusing finetuned models for better pretraining.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Fusing finetuned models for better pretraining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.408331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.408331Z digest=sha256:cb9017fc87ef2c492f641d81f99ff715759d004bf54f58502e3fab995094877d

Observation 60b0b9d2-5d8d-402c-a46a-6a79e9e9d94b · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.412285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.412285Z digest=sha256:c28c8da50247575b3133325290721dfe76fecc6c49ea6b34b235d96071e423f2

Observation 73f5ab54-1ee5-447a-8adc-e30fed190887 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.416519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.416519Z digest=sha256:a912c01eaea983351f8642a10b39c8aa4f13235a5e9192e23a09f6a8965c23e0

Observation d5fca2ec-fa38-4895-bcc6-21d032854460 · outbound

This paper cites Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.086718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.420385Z digest=sha256:374591863c206e37faf50ffc23a2bf835624cdbdf9c6842360eb384619a7c422

Observation e84ceb55-0f61-4ce1-b195-7ba48fdcab56 · outbound

This paper cites Understanding Emergent Abilities of Language Models from the Loss Perspective.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.423803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.423803Z digest=sha256:1d35bd9bacd74c6c5a2f556de32fe5251e76bd09756739f68eae180f7016b35b

Observation 78e44852-42e8-403c-b09c-12ae7f0cb577 · outbound

This paper cites The Llama 3 Herd of Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.427749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.427749Z digest=sha256:2efc5a8bd73f8946262530da1d915db0d737b9fd319ecdc5d318791b46217f5c

Observation 253c7a5c-2696-4055-82b8-a1b0bab31ecc · outbound

This paper cites DoGE: Domain Reweighting with Generalization Estimation.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoGE: Domain Reweighting with Generalization Estimation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.431339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.431339Z digest=sha256:6b697f6ca8bca5d37661b2a66e334148b34b5c10613a3be33e3315a7ae6264ae

Observation 161d34cb-2394-49f7-9871-818fd68db1ec · outbound

This paper cites Dynamic Gradient Alignment for Online Data Mixing.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dynamic Gradient Alignment for Online Data Mixing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.434779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.434779Z digest=sha256:fb463d2b895d8c1985f7a8d89d4a3398b08753e9f637c22080fa1d464da4b295

Observation 11e84925-235f-4796-9d8d-0207ad581f70 · outbound

This paper cites A Review of Sparse Expert Models in Deep Learning.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging A Review of Sparse Expert Models in Deep Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.438208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.438208Z digest=sha256:c4fe6dea799cffd4370ba77b364367becd2cb6c1ab1fa964c586112a9cdd36aa

Observation dfe26220-49f0-4b71-a9b3-2ebac4de105f · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.071793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.441375Z digest=sha256:e734d97f41aeb5cf709b50e8c68a66ed69b853e81bea65585108ab54f21ef956

Observation 51ea79ca-b8ea-4326-9a21-2ceaa9d89f8f · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.445085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.445085Z digest=sha256:d32cf81b8db698975840172999c68cc6f552209e03fb1a17bfb675b8ec338fb4

Observation e56c246c-31cd-4f83-a061-a65fcaf2b04a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.449094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.449094Z digest=sha256:0cbba697553e849bab7797e8cf363ba1ab3866ab05caa237eda904f8f72da34c

Observation c269444f-8b4a-4398-a1f9-52d7e5fa46c4 · outbound

This paper cites Demystifying Prompts in Language Models via Perplexity Estimation.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Demystifying Prompts in Language Models via Perplexity Estimation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.452868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.452868Z digest=sha256:81fd6e3a25800622e4dac84450834ff2c742b414e9c82300df44fa555e9ab7a1

Observation 82923042-9307-4469-b0eb-b2f623bfbf6a · outbound

This paper cites Adaptive training distributions with scalable online bilevel optimization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Adaptive training distributions with scalable online bilevel optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.059125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.456462Z digest=sha256:ec3d70079e85511166c59e1290b3799583eed83ae07d7bf5100fab85a4140b89

Observation b08e76ba-e5e0-4789-ae2f-033f557d9885 · outbound

This paper cites Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.460128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.460128Z digest=sha256:edc6a80530f3133883d7a4d85122e441055206adb26c3ba4f3c8522bf688457d

Observation 84c69eca-4bc3-4985-9bb4-c61d6c11baf3 · outbound

This paper cites Hard mixtures of experts for large scale weakly supervised vision.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Hard mixtures of experts for large scale weakly supervised vision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.046874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.463977Z digest=sha256:b5c8e8b5abf0e3aa96d96adb3decb4b457eb1603cb8325b1b2422a42976679bd

Observation 88c8c2b4-1ca4-47fa-8016-cb09b1ebf898 · outbound

This paper cites MiniLLM : Knowledge distillation of large language models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging MiniLLM : Knowledge distillation of large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.034269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.467487Z digest=sha256:5a76fe3faca0400cbd1ead5792612091b5800196775eb2e967e15161db58d0ae

Observation d14bff1a-2729-4624-a010-25750b77502a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LoRA: Low-Rank Adaptation of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.471035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.471035Z digest=sha256:2ed960b99090706120544b535a6de097e8aa7476a919348d44bffed51a5bab73

Observation b40ae42d-0f3e-4bb2-b432-76a066acc317 · outbound

This paper cites LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.474822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.474822Z digest=sha256:1d0eb99910f15c79f681f489f6b237e4eb9778012b838c32b431551acb438c56

Observation 0c026c98-4752-4671-8734-56aaf4aeb769 · outbound

This paper cites Editing Models with Task Arithmetic.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Editing Models with Task Arithmetic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.478584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.478584Z digest=sha256:8058362c61b39d69add5070291071825151142cc37ab4c7f72c237f22ed2ecbf

Observation b3d6253a-fe63-4121-ad08-89e0ec34babb · outbound

This paper cites Mistral 7B.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Mistral 7B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.482212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.482212Z digest=sha256:aae726847a78f7868afe362801c73841a4fe32ce880008485db9a7ecad78d649

Observation dfde294a-c936-4492-9067-f943d5e6baa1 · outbound

This paper cites Mixtral of Experts.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.486070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.486070Z digest=sha256:08f059eeb08f6a278bcf01137b6ca589c27dd7566f7194d0097ab4f8511168a2

Observation c19c0c4b-8781-48de-b369-9a39376106a6 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Adam: A Method for Stochastic Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.490125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.490125Z digest=sha256:16e9cf6b45b860ae2e2d7a8624905382ea1d868c48e394d5ce632fdd10cbcbd6

Observation 2745ea1f-7427-405e-851a-79a3be989fa8 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Scaling Laws for Fine-Grained Mixture of Experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.493957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.493957Z digest=sha256:e9b6110ff626858d97075957d19fd0e5c8fb383cae30862472f0acb04b2688ab

Observation bf38f286-0ecc-4312-a1c4-35e60eafe9e7 · outbound

This paper cites Evaluating quantized large language models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Evaluating quantized large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.021236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.497889Z digest=sha256:0c212d4f98122368808df54758970e0b54e43a53e7e29d3a874a71f36c9292c3

Observation dfa78e05-8bbf-47bd-9645-58cc9e40dd9a · outbound

This paper cites LLM-Pruner : On the structural pruning of large language models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LLM-Pruner : On the structural pruning of large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.007734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.501526Z digest=sha256:265d0d74fed4d9ba69fbbb7b1eac725104014586a2a085acdd6bc7af566f6240

Observation 58b1f7da-f0d7-44b4-91dd-d26fec265c93 · outbound

This paper cites Task arithmetic in the tangent space: Improved editing of pre-trained models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Task arithmetic in the tangent space: Improved editing of pre-trained models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.994722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.505793Z digest=sha256:f81bb8026b7280a4541d9e80c4603b5fffce030693da956ec0389931b10b02e2

Observation 427b3fff-99d2-4f59-8d94-36d76c0383d9 · outbound

This paper cites Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.509397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.509397Z digest=sha256:657c5a32a66bac61884476933157a2f1abe8e7fc6b6525d2d548644d15dab07f

Observation 4f337655-4d5b-4423-9baa-1ac280c93dfa · outbound

This paper cites Diverse weight averaging for out-of-distribution generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Diverse weight averaging for out-of-distribution generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.981377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.513283Z digest=sha256:961d32ecedd18b37fee3efd5dcdd0ff89f6326694a7624f05c6844d3720273eb

Observation 46644d6d-fc48-48b0-bc87-de9253bcb8f6 · outbound

This paper cites Model ratatouille: Recycling diverse models for out-of-distribution generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Model ratatouille: Recycling diverse models for out-of-distribution generalization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.967482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.516804Z digest=sha256:517bd4808361175a6001c3ae1ffe0469ed5204b3f4f7c918512c36a5a668352c

Observation 27462a9f-c869-4aaf-9e3c-930437dfe036 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.956064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.520952Z digest=sha256:d1df8d961a37aadfea784c94abbf43b6bcf0024184bb02d0dd935c80e79c5321

Observation c4d32ccf-b590-4ca8-b125-6fa924230722 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Gemma 2: Improving Open Language Models at a Practical Size

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.524414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.524414Z digest=sha256:a2f553458d42d8e0d395d0bd67e044f4ff7a25c754f49c8d9e68cef5cfdab53a

Observation e4f29297-d2e5-435b-b527-e50f9296ed6a · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.527885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.527885Z digest=sha256:230f12de28cfe30f5d0d4c343e6f6627b8796a5c459bcc3082ceea25ee41d76d

Observation bdf10fbd-1173-4d8d-92af-32373de6e98e · outbound

This paper cites Realistic Evaluation of Model Merging for Compositional Generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Realistic Evaluation of Model Merging for Compositional Generalization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.531857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.531857Z digest=sha256:df52c6774a5235e427d35b829414701f540cdb16414ae5c3e5b3b78299bd849b

Observation 025f7d6c-2553-4731-87aa-2563046a66ed · outbound

This paper cites Efficient large language models: A survey.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Efficient large language models: A survey

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.944220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.535349Z digest=sha256:66e71f975ee49f4478a45e10a5f676b40205cb3785b294379245d215834ddd42

Observation 3043a95d-cf58-44b5-9822-d5c72ff727a3 · outbound

This paper cites T., Wu, T., Song, D., Mittal, P., and Jia, R.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging T., Wu, T., Song, D., Mittal, P., and Jia, R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.931671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.538832Z digest=sha256:8253d6c03353fbc9ce81e02657887660296f1061430301caf3bed28a17ecb185

Observation 2faa9b3a-7af2-45e2-becb-87d3f888bdad · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging RedPajama: an Open Dataset for Training Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.542454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.542454Z digest=sha256:e71fbb46520c2b0fa38f78619ada986cc702137c9e5b0f366c870713ef191b02

Observation 4d0d2575-1cf5-4df2-a85a-ac6c08caa8c3 · outbound

This paper cites Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.919199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.546212Z digest=sha256:7cb64beeee04c215d8c41b22397cfa0cf58fe79df7ed6b180804347ae1cf0eb6

Observation cda5a5a6-bfab-46eb-a2c2-8ed6b42dd7dd · outbound

This paper cites Structured pruning learns compact and accurate models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Structured pruning learns compact and accurate models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.549919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.549919Z digest=sha256:6efcd0e8d3e2ec69729ce828c22079b554d31f79b6ad787396382793d6f50fee

Observation bb2ab740-b7eb-4d1f-ba23-f443ee314bdf · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.553441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.553441Z digest=sha256:c50f280828780a3dc401e89c9aca33becbfe1cc5b1b8b30da9929abe9babf0eb

Observation aa434cf1-ebd4-4aef-9ead-0802752e0920 · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.557385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.557385Z digest=sha256:ee63dde86f5ed99d10daebc7a25634380a36b447df2192b9321f5ed1fe30fd56

Observation c1156ac8-7ef6-4773-9ce7-1b3eacfd4c53 · outbound

This paper cites write newline.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.561399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.561399Z digest=sha256:5f8637321c38aca49956da291f96a2a1300f2d24de99ce0f34d9d40e6031b1ad

Pith citing papers

No inbound Pith citation observations are available.