Pith. sign in

Paper Citation Record · LEDGER

Upcycling Large Language Models into Mixture of Experts

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.07524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.07524 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:21:00.725253Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:07:09.402995Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 08b928ef-fcaa-4d83-8a1b-653c58669297 · inbound

Scaling Laws for Upcycling Mixture-of-Experts Language Models cites this paper.

Scaling Laws for Upcycling Mixture-of-Experts Language Models Upcycling Large Language Models into Mixture of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:00.725253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:21:00.725253Z digest=sha256:2606020e0c3b0746a7617bfb78242657291f7d13b949427d9f8305636b5246c9

Observation 429f91ef-ec6b-419b-85c8-820860df9c47 · inbound

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights cites this paper.

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights Upcycling Large Language Models into Mixture of Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:58.896820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:58.896820Z digest=sha256:360534515b8d1c477581b3f486ccd2075010cdf8cf67381f1d1837aa7011084d

Observation ca0b421a-e734-4678-a90d-d4c2e5a5216b · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Upcycling Large Language Models into Mixture of Experts

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:05:47.640835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:aeb5711a0b3eb162cf89ffc9cdffcb0e4b6439ccf63c2c3b1f90a73cd88109a8

Observation f319c1af-7c8e-430a-851d-df31903dabd4 · inbound

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts cites this paper.

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts Upcycling Large Language Models into Mixture of Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:26.187530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:56:26.187530Z digest=sha256:594070af7bcac005b5084bae70af6a5f0d41bca55daa0bb5b5666cd14634854e

Observation 25135784-4b59-4a67-972e-feebfa6fbd7c · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Upcycling Large Language Models into Mixture of Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:51:45.742897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:8c030d35e5923fb1182e7403007832fc2fabd5efc134176a108b83959cfcbb34

Observation 948fbb49-4471-4e45-9fa2-e49c6de2f255 · inbound

Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts cites this paper.

Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts Upcycling Large Language Models into Mixture of Experts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.469330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T20:36:22.974054Z digest=sha256:c285a4556d4980acaa6759a24d3fff77b6668520518bad250bfed1a434189221

Observation a7a11777-3a83-4b70-9ae5-76f91184af70 · inbound

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics cites this paper.

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics Upcycling Large Language Models into Mixture of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:05.261852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:05.261852Z digest=sha256:b5a64650ae0cfb8b5d9beddf98fa94765d93d6b27cd2b44cb750260b6eb165e0

Observation 10e95024-95ad-42d7-80fa-9e0311fbef25 · inbound

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics cites this paper.

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics Upcycling Large Language Models into Mixture of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-15T11:44:19.622453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:44:19.622453Z digest=sha256:87f1013df6bc7687820b82430450381907c12a31023988c51c72a8ea28e8dd5d

Observation 233874d5-5b0d-4533-9c9a-2a547347cac2 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Upcycling Large Language Models into Mixture of Experts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.490090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:29:16.555166Z digest=sha256:e69899832d11f49535fcca484eb58d983e7983feab56a167171e20d50082bb1f

Observation 2369a1ef-cacc-4879-b676-cfe75af1b087 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Upcycling Large Language Models into Mixture of Experts

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:15.333831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:03:02.654035Z digest=sha256:12b67db345524eef09db86ae7123a0d33e6193835d937849ed4d27f44b72e4c6

Observation 6eaade6c-b6fc-4c73-a5c1-80b6ce4ccb6d · inbound

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey cites this paper.

Tackling Multimodal Learning Challenges with Mixture-of-Expert: A Survey Upcycling Large Language Models into Mixture of Experts

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.302417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:26:44.904121Z digest=sha256:049a236de2c7c2f274b7f30eedbe5c2676adcb7e4533043b9d0a4cca3ef76e00

Observation 27f8962d-9ae7-4be6-8894-b4aba9fbcf7e · inbound

Reversible Foundations: Training a 120B Sparse MoE through State-Preserving Scaling cites this paper.

Reversible Foundations: Training a 120B Sparse MoE through State-Preserving Scaling Upcycling Large Language Models into Mixture of Experts

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:07:09.404976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:55:09.477413Z digest=sha256:badb8c67bcc145451effef6abae5dba41bfb41086181c1eb249840995551e802

Observation e16c331f-e8ef-4bce-afb7-c56184c45523 · inbound

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition cites this paper.

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition Upcycling Large Language Models into Mixture of Experts

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:46.956823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:18:46.956823Z digest=sha256:93b1459c812b720851b4e7a1c406a35867e92daa01f9c02c11680b5be8fedd0f