Pith. sign in

Paper Citation Record · LEDGER

BASE Layers: Simplifying Training of Large, Sparse Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2103.16716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.16716 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:29:49.048895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:03:17.812816Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 422d557e-f2d9-4302-9520-2311b527c97e · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models BASE Layers: Simplifying Training of Large, Sparse Models

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T23:14:25.641570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:985e462c76f0ff3591a4bf52c032d35c8c16b8c47a4cf8e45132a6e587a84c61

Observation 016edb7f-bd71-4238-9d6b-453ca3201168 · inbound

OPT: Open Pre-trained Transformer Language Models cites this paper.

OPT: Open Pre-trained Transformer Language Models BASE Layers: Simplifying Training of Large, Sparse Models

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:53:17.632763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T20:53:16.720145Z digest=sha256:ffe1c070bca4cdf1f27ca09092427563606e63c709334a53af1cbbe122acb728

Observation fb0dd2ee-7687-41a9-b492-3a32e9295655 · inbound

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale cites this paper.

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale BASE Layers: Simplifying Training of Large, Sparse Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:35:36.151655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T13:35:35.972596Z digest=sha256:91bb32481a8c071830461e76309e42b0d586412d265be223e110b2b69b2ef9ca

Observation 66d12c9b-7f6f-4926-8994-39f71ca8ff5a · inbound

A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router cites this paper.

A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router BASE Layers: Simplifying Training of Large, Sparse Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:03:17.814167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T09:28:02.669459Z digest=sha256:92da5be0921bb4f0317d51d75d9f4611bca46a75eef91c7628f294a62ecc05ac

Observation a98cf91f-94b7-4bdc-afd9-d0719badc588 · inbound

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism cites this paper.

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism BASE Layers: Simplifying Training of Large, Sparse Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:58.299078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:58.299078Z digest=sha256:99aef885da5caa6f381702a53b30b7e8a9f42a072cfa441753d9d428d0ee359f

Observation d6a0a8a9-c102-4e9d-9621-8e5174d5fa11 · inbound

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry cites this paper.

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry BASE Layers: Simplifying Training of Large, Sparse Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:29:49.048895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:29:49.048895Z digest=sha256:6787fd1c77e7223aae3704c53ad8a54997b88dfae9797e0fede2e84ef6dedee8