Pith. sign in

Paper Citation Record · LEDGER

GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2112.06905.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.06905 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:09:26.809525Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

168
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f6ac1ab7-2b5c-4907-ab43-e42737955352 · inbound

PaLM: Scaling Language Modeling with Pathways cites this paper.

PaLM: Scaling Language Modeling with Pathways GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:07.185970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T23:45:06.755839Z digest=sha256:89356fa957a83ae377fcff14359494902b7386b2429ae4c2a324259f308cb41f

Observation 40bc1e35-4afd-4601-a92a-77bdfc24d024 · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:38:38.291396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:900d74d4f2ccaf3f8a4a4cc59d031afbb8b41014fbf7f5fe557389f9c6ba365a

Observation 545b5eed-f442-4098-ae5f-386a3b60a53f · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.100632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:844b106d4d4239e82c0e8e5bda194e3b773bcd9e52443a6ab49f1c4cfa3d6ecd

Observation 3d053dd6-2f60-46aa-928b-b0d7114fc72d · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:41.895334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:495f0a4debee96f5f5e65f6d364e27155896d5b8b64e04e658b4e10e542bd02a

Observation 470adb7a-1458-442b-8356-efc98aa84487 · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T11:59:27.202634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:37ed13926b68211a91afff4adf47d93a4b5fde1fa606dd126929ed7798584cdf

Observation 9fcfa9f6-ed31-42d1-86f3-3925a5abbb97 · inbound

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models cites this paper.

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:22.297557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:07:22.166595Z digest=sha256:0a7dc1f6794d0ff0ead0617635da4d3a436dadee4577d4ca0a0916351ae2b517

Observation 750675db-14cd-4cb7-83bb-84120156c793 · inbound

CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling cites this paper.

CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T17:09:26.809525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:09:26.809525Z digest=sha256:51422e15842b6a16886c339d8949e4c025ee25b6c3b2498699da47543e8ab064

Observation f6e487e9-a792-4216-8855-6e30d3dee94a · inbound

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference cites this paper.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.440954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.440954Z digest=sha256:374ea653bb85c7c2533983ea81b979e3b752409bc45200ab6177034aa63822cd

Observation 963250a7-090d-4ab7-9b92-d715f9f2a4b9 · inbound

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives cites this paper.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.759942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.759942Z digest=sha256:4ab8297b4549a7e54ad50a7547209f7916d2f7a250b5c60834e42228c758e707

Observation ac690ccf-7837-4ab3-86b4-91ad969c3ef0 · inbound

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights cites this paper.

Scaling Fine-Grained MoE Beyond 50B Parameters: Empirical Evaluation and Practical Insights GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:58.233961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:58.233961Z digest=sha256:9788cebf702b5cf125421e74fa209e977cbc048f9412e0ab21062c4eddec5215

Observation d1bb6067-0114-418b-88a9-1221a291043c · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.175386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.175386Z digest=sha256:e25af8b44e4de4d488abc88b21ae8e84c3e88cde5e4e265fa6c84c541a3f77d7

Observation 35e6ba7b-1919-4fa1-9b66-094c5418f6e7 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.921973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:042babf00e239a2bb2df18afe0a49e9596f6d6ba1565074ceda891f47e4ca3a3

Observation 33aae63d-db23-4645-aa10-691cc4d1adb6 · inbound

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling cites this paper.

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.180698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.180698Z digest=sha256:bbb5b06c60a03801dc9f35bdfcc1859eb5653dd41401d55ced646c900d292626

Observation 3cef134b-8794-431d-9533-d59c13455cff · inbound

Apple Intelligence Foundation Language Models: Tech Report 2025 cites this paper.

Apple Intelligence Foundation Language Models: Tech Report 2025 GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:58.562199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:58.562199Z digest=sha256:5f2f76b2e06321450f88949da2f0273b4793a15a079c3242388cc34cf9e0cc52

Observation 0f324ac6-b139-4f9d-ad51-d0c8b51e4080 · inbound

The Carbon Cost of Conversation, Sustainability in the Age of Language Models cites this paper.

The Carbon Cost of Conversation, Sustainability in the Age of Language Models GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:13.238051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:13.238051Z digest=sha256:440c51269cabad8f0ad6b05157f3ee91669b5b969f0b70015fb45091e16aadcf

Observation dff329e2-b034-45f8-86b2-a43ad82782b3 · inbound

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned cites this paper.

Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:28.421116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:28.421116Z digest=sha256:b3a170201075ede1056d150d2d30da3dadff7a47f3f0c14a0dea4c73916eebd9

Observation 8cf3ad8a-9b48-4d79-aaee-355c18ad0f5c · inbound

Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression cites this paper.

Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:26.916040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:26.916040Z digest=sha256:fe851cfaa3eada236e62f844ec873d768f19fefe97cfb00f42aad13833caca88

Observation 114051dc-82fc-4956-ad02-29751a4ec932 · inbound

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents cites this paper.

A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:24.535811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:24.535811Z digest=sha256:d75bf9289d83b97674e97f0193707b5d6b102f85ef8d7bf02581fd69b99ed3f3

Observation 9a54603f-628a-4d12-934c-30c4f2595c77 · inbound

The Ray Tracing Sampler: Bayesian Sampling of Neural Networks for Everyone cites this paper.

The Ray Tracing Sampler: Bayesian Sampling of Neural Networks for Everyone GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:36:36.566867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:36:36.566867Z digest=sha256:6b17d302629f73e05a1251a4166c634a8f42997683bae94f1b7ee8787b5a3ca7

Observation d97d9803-7b68-4eac-afc3-3642c2f172cd · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.519092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:29:16.555166Z digest=sha256:505b280d5684c4aba3770f009eea7cd46ec605ea031dd621966ea122836211fc

Observation 02157ec7-66e4-4b20-9e39-19a3021f6297 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:15.356374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:03:02.654035Z digest=sha256:7bffe50620f5030aa80a045d244b04b1033337dee7fac8233796ca06c8b327f0

Observation 36039638-87a5-4c8d-bc05-45dd6f603b3e · inbound

A Meta Reinforcement Learning Approach to Goals-Based Wealth Management cites this paper.

A Meta Reinforcement Learning Approach to Goals-Based Wealth Management GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 262

Resolution
verified exact
arxiv_id, observed 2026-05-08T18:44:01.879605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T18:42:50.962120Z digest=sha256:28ba624918b975a65af48418a80ecdb2e507fffbade37a40c9cc4722bf4228c1

Observation a53b1ea1-3b78-4027-ad55-e6f0b6507373 · inbound

ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems cites this paper.

ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.565438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:26:14.798902Z digest=sha256:fae15b173e2a6061fc82af8330e3403edc2161595a453b3691abc934e8bb849a

Observation 24775fef-9c92-449c-acfc-60660f40295c · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.058669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:11c613d5801322cd78dba8b85c30d7cc6f99d77050e61d4bd74b3b25e3520e18

Observation 3e2be5df-f02b-4fcb-84cb-790b57ddc27c · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.809933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:9c3498805f16bf51fa514b5c5bc17608532bd67d359315d1ab09e7fa11b48925

Observation 67c4ef99-73a8-42ff-b43a-b2b7265ea495 · inbound

Memory for Large Language Models cites this paper.

Memory for Large Language Models GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:37:54.336046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:37:54.336046Z digest=sha256:4892f94f9226a3d5706c2bcf792c63d8522ccea1207945dd17a768f415f8b39f