Pith. sign in

Paper Citation Record · LEDGER

Efficient Large Scale Language Modeling with Mixtures of Experts

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2112.10684.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10684 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:33:22.720381Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:36:16.721505Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32a6ba29-5386-4c8e-a5f7-0db809e90907 · inbound

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? cites this paper.

Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:51:46.877478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T09:51:46.701149Z digest=sha256:1bac8cd6c9f22f163504bdeb8ea44c147b678fa74f9f23b6baef3c1f7cdf8b67

Observation f55a99eb-5a16-4349-8d7c-42307fc0760b · inbound

InCoder: A Generative Model for Code Infilling and Synthesis cites this paper.

InCoder: A Generative Model for Code Infilling and Synthesis Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:21:20.464819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T02:21:20.438666Z digest=sha256:3cc3d80a464b00e93f16dd71a9f611749f7aee94f22ceaf161d8255c4552291d

Observation 87fc9f96-7a52-45ad-a387-688048f16105 · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:34:28.200651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:474f3d8aed2f2d3ab0cf67b697aa9f58d00daa7181efaeafb275366294c80347

Observation 1cb42d52-0636-42b9-97ed-e67f246027a3 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 287

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T20:53:17.887823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:1dda66e72b78d350463d7c1e033b535dd42f7e3df611e318ba77f378b43feb23

Observation f487086e-f9af-4a06-a760-1fba15f9477d · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:38:38.018560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:ed28706c88e66e8edd943e1a40ba308ceebaa2805160e565d3ed80f98ef9cc47

Observation 33f09b31-2c0a-4d40-8320-75c228c7512c · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:35:36.034981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:cc27b557bf2d8ed3423fbd88d1a5da1558200d73b36d60a88b2cc428bd5753cb

Observation 4cfcdc35-95fd-4e6d-9bea-35bfe4e126f3 · inbound

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers cites this paper.

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:44:22.789288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:44:22.710205Z digest=sha256:9ab829ca21c267c110b14587dbd0b96ca9988e4c59d4c6c336604898ac86bb14

Observation 9582cd58-99b0-4787-bf69-af2e15d3b682 · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:43:45.814699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:b58f6fd6fe4728e444882ff59fa49b5af8c2c3032659ebd238501de5958d53ca

Observation de023224-8b76-4114-9561-e053f1dcdced · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:22.300478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:42f9ae84bfd654bd358130f5ab392d1c247b18ee56ed0ffeba516f06450b07c6

Observation 4209e00a-e48b-4842-8231-8b05048a10d6 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 242

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:09.982830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:37c9828b7d7ee548e245f2b4b13e28073539ee28053f53dd13d1acbda4ec8bb8

Observation f3dd5bc0-b804-493f-999d-4fe705cc09d8 · inbound

Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction cites this paper.

Faster Machine Translation Ensembling with Reinforcement Learning and Competitive Correction Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T14:33:22.720381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:33:22.720381Z digest=sha256:f2bcad67bd3cfb7a83057587a530497562b792804bac41f2553b0193a4eee904

Observation 0bcc0a66-3e83-47ef-b5f1-7647ffb53a27 · inbound

A Survey of LLM $\times$ DATA cites this paper.

A Survey of LLM $\times$ DATA Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:12.148766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:12.148766Z digest=sha256:87131525ea74c53663d9b216a3be772b3beb761425d76af982dc8cd5e5e3fe62

Observation a320d8a0-be26-425e-b3e8-683d7ac95a5d · inbound

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO cites this paper.

A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:52.178001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:52.178001Z digest=sha256:7e4d79bb96982962d1c0ab107db29f9f55b5a0201e49c4989d231928ffd29ec4

Observation 47e4f7dc-5113-4cc3-b083-c681add78bde · inbound

On the Fitness Landscape in the $NK$ Model cites this paper.

On the Fitness Landscape in the $NK$ Model Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:28:53.637326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:28:53.637326Z digest=sha256:0361216057e62f2092a5ed7ccc2c7123d10725090ea27845063b443843ce84db

Observation 78aa2a76-a5c6-4126-907b-2a602040a306 · inbound

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference cites this paper.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.504969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.504969Z digest=sha256:5b251795e8d8ea46493d28ec2a622ff62732ad52e522054c7e9410033e81206c

Observation f403ef49-7024-4c19-a63d-0f28c93c1b59 · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.224496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:014168015dbf7bbcd254fa7a0098cf11594197a00d6caab8e0ac3741434061ce

Observation af5eecf4-66be-463d-a200-6a4e94ef4ba4 · inbound

Tracing the ongoing emergence of human-like reasoning in Large Language Models cites this paper.

Tracing the ongoing emergence of human-like reasoning in Large Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:04:37.198324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:04:30.829386Z digest=sha256:cab05c12f1d0ed9cdd2b04bc438df7246a072fbc7b98fcbd92a837fece19e205

Observation 54776406-cfa2-4997-b3f1-36e6e72426e5 · inbound

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs cites this paper.

Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.034915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:22:40.017922Z digest=sha256:783bd8ec00534eb11ebc2f518d461e030462852a7e48762377f02710ca18ab41

Observation 072673fc-84d4-4fd9-8b66-1fd0f9e22efa · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.722971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:8c6dd3747fe02f169b1714f6257d9c263579a937d9dff50faed23ff7046ca75e