Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Scale Language Modeling with Mixtures of Experts

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2112.10684.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10684 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:57:52.538168Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:16.721505Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32a6ba29-5386-4c8e-a5f7-0db809e90907 · inbound

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? cites this paper.

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:51:46.877478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-15T09:51:46.701149Z digest=sha256:246beb21598079a82abe8310ef0cd1f44a72720898f7bf38982668db19b6abd7

Observation f55a99eb-5a16-4349-8d7c-42307fc0760b · inbound

InCoder: A Generative Model for Code Infilling and Synthesis cites this paper.

InCoder: A Generative Model for Code Infilling and Synthesis Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:21:20.464819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T02:21:20.438666Z digest=sha256:59a1cc2f85b81cae48ff5d4bde20173255b9ae4a79791790faaeb1431e41ca76

Observation 87fc9f96-7a52-45ad-a387-688048f16105 · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:34:28.200651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:6dbbf085a46d1ff54e71d2bef9cff1ada36b354320a0ae909b8205cf62ec1f0a

Observation 1cb42d52-0636-42b9-97ed-e67f246027a3 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 287

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T20:53:17.887823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:124b81cd8bd24363833a3f9302bc54c6c43587c3954c0cf77c6ff0dbbb098205

Observation f487086e-f9af-4a06-a760-1fba15f9477d · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:38:38.018560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:87738a919b4274fdcd4123a34d89da1ee44a7192c8135c129e2fc285b5e75bf3

Observation 33f09b31-2c0a-4d40-8320-75c228c7512c · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:35:36.034981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:025074c773f5e3e8414133c3a9b29d3ff0be50a419b38a9f4eafa4e2cc9dd524

Observation 4cfcdc35-95fd-4e6d-9bea-35bfe4e126f3 · inbound

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers cites this paper.

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:44:22.789288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T01:44:22.710205Z digest=sha256:559e80aa97809b74fcea8de13329b5570e8cd1d07a6cad5a171c54c72a449cc0

Observation 9582cd58-99b0-4787-bf69-af2e15d3b682 · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:43:45.814699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:a92031ed5842133d9e74f5c0002ded3f9cb95d321faa1f5ad2db7c68b95bf2f1

Observation de023224-8b76-4114-9561-e053f1dcdced · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:22.300478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:b086b25f17f02d23fc45c851929cda0a92c8721c2244fd496457e3005ca8f251

Observation 4209e00a-e48b-4842-8231-8b05048a10d6 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 242

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:09.982830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:38ee6345a24ca52ff8589a0c0c8f2edb79238a8228c4bee06e658973d995771f

Observation 63cd8aad-358f-4533-9ea0-ee24ee7c7a65 · inbound

LLM4CVE: Enabling Iterative Automated Vulnerability Repair with Large Language Models cites this paper.

LLM4CVE: Enabling Iterative Automated Vulnerability Repair with Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:52.538168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:52.538168Z digest=sha256:5f4093ce2f9295d188f0a6c225a30af1bf575c27edee483f51e6ee78a611f769

Observation f3dd5bc0-b804-493f-999d-4fe705cc09d8 · inbound

Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction cites this paper.

Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T14:33:22.720381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:33:22.720381Z digest=sha256:e0bfe5b36005ad62686d6a76f6016df28278e20aec65810aeda3e3628ad21240

Observation 0bcc0a66-3e83-47ef-b5f1-7647ffb53a27 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:12.148766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:12.148766Z digest=sha256:87131525ea74c53663d9b216a3be772b3beb761425d76af982dc8cd5e5e3fe62

Observation a320d8a0-be26-425e-b3e8-683d7ac95a5d · inbound

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO cites this paper.

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:52.178001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:52.178001Z digest=sha256:65b59fd06434f362270e2e6ee8a100e4d29fdd2d3ac9401615387f3541e97c6a

Observation 47e4f7dc-5113-4cc3-b083-c681add78bde · inbound

On the Fitness Landscape in the $NK$ Model cites this paper.

On the Fitness Landscape in the $NK$ Model Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:28:53.637326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:28:53.637326Z digest=sha256:c37a94a4768776d6937ea8381f645fd35688c767df7dc5314808218901666611

Observation 78aa2a76-a5c6-4126-907b-2a602040a306 · inbound

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference cites this paper.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.504969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.504969Z digest=sha256:5b251795e8d8ea46493d28ec2a622ff62732ad52e522054c7e9410033e81206c

Observation f403ef49-7024-4c19-a63d-0f28c93c1b59 · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.224496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:012d9a4b7c17f0d8f10f6e8294d67b7355519ce1e2a3c116bf2e3585ac5f54bd

Observation af5eecf4-66be-463d-a200-6a4e94ef4ba4 · inbound

Tracing the ongoing emergence of human-like reasoning in Large Language Models cites this paper.

Tracing the ongoing emergence of human-like reasoning in Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:04:37.198324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T05:04:30.829386Z digest=sha256:0876e8a726ad9705a8687937171c2b77b15ec19e88c70ce3bed918f67a2a0f01

Observation 54776406-cfa2-4997-b3f1-36e6e72426e5 · inbound

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs cites this paper.

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.034915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T13:22:40.017922Z digest=sha256:201f883ece1804fef395e9ef5b2ddc0e8444faa73f14a03de715f7c45bf8a714

Observation 072673fc-84d4-4fd9-8b66-1fd0f9e22efa · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.722971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:ababb57a0259687304f2cf5920c898c4ec1bfaa7e28807f63006df0ba49cf829