Pith. sign in

Paper Citation Record · LEDGER

Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2304.03208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.03208 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.474327Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 72709979-5cf3-409f-93c8-87bf672b1828 · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:43:45.805619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:93ab9b0028435f74707778ff9365897ea9bc66e7035fe3c85fceacf9433b3c90

Observation 11ce8e20-346e-4ee5-9948-e01920c03746 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 273

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.073190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:fb2ad4607b441f7d8031ca86bec744102248e6dbcc5ec4db1fe9e9360a959bb7

Observation 7ded08bb-9350-4ace-a73c-fa0057af11eb · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:00:53.450722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:dff1b4aa62331f338d69f315d7e6d9eff51b622b315144074a1f799af69f1cee

Observation 7f2cbd0f-da97-422e-8f01-f1548ff192de · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.423436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:48281d6bb95a22d5567e22d54a0bee6136d62bcb1ab977cda3ae1fe3a37d073d

Observation 8d69a650-a8bd-4bfe-bada-f1598e01d261 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.474327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.474327Z digest=sha256:10c1e4dd596658ae36a6831067720221812e8190d3d09e43e47c773ab330f918

Observation 0d10b42b-515c-40a5-b47a-0fafba6430f3 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:23.116345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:23.116345Z digest=sha256:87275c964d4d9f424f9431d5244e3385c5cd18644d6e08a560ff21301617ff91

Observation df1ccc47-c817-46e5-b64c-4ed0e368cc4e · inbound

MuLoCo: Muon is a practical inner optimizer for DiLoCo cites this paper.

MuLoCo: Muon is a practical inner optimizer for DiLoCo Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.558466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.558466Z digest=sha256:d12e49f9daeab3b4b29aa95f446413fa2674b4f658c5de5ad71314035386707b

Observation 7b19a67a-b969-4e9b-bd05-f3e374efaaa7 · inbound

Basis Transformers for Multi-Task Tabular Regression cites this paper.

Basis Transformers for Multi-Task Tabular Regression Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:18.794719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:18.794719Z digest=sha256:d4a89d36476d64ec58fe6c4d982a93ee5a2292a04202de768f02535295540050

Observation 4c52a78a-a6e0-4ad8-a59b-a3e36ad36ad0 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.928577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.928577Z digest=sha256:9d45fac17de67e925642670ef33aa91e4641ba4f0cfed7563dde5969d727e4f4

Observation 696b93b9-96f5-49dc-af04-c4a478f741f8 · inbound

Similarity Field Theory: A Mathematical Framework for Intelligence cites this paper.

Similarity Field Theory: A Mathematical Framework for Intelligence Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:41:30.649201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:36:49.543497Z digest=sha256:e56487e50254a0cca414f0ef81b9815a17a9ee1cce9a37c3ef000f5f396c1844

Observation f6d2d069-c6f6-451d-9f9d-23f49452f5d0 · inbound

SpaDA: A Spatial Dataflow Architecture Programming Language cites this paper.

SpaDA: A Spatial Dataflow Architecture Programming Language Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:20:23.248975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:18:38.464461Z digest=sha256:dc6a04777f3ade63ef93e1dcbb28d075a8ceb0b390b9aacc65df23d106adaedb

Observation 1c9812f9-a439-4926-ad6d-d58639597ea3 · inbound

Spectral Condition for $\mu$P under Width-Depth Scaling cites this paper.

Spectral Condition for $\mu$P under Width-Depth Scaling Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:06:25.514696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:03:31.202134Z digest=sha256:0559eed8e738d4c97b518f1f0c2f18898d8e5fe56c9a2668b1dd293fd0e1aa65

Observation 3030597b-90d3-49d4-b542-8bcf4d65b7e0 · inbound

Learning the Signature of Memorization in Autoregressive Language Models cites this paper.

Learning the Signature of Memorization in Autoregressive Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.472570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:53:10.396785Z digest=sha256:fb276a0fecbce35d90c48b4810e101b699ce1e50b94b79af7cc4b8d939dcdb83

Observation 397fe58b-91bd-4d8c-832d-adabb3ac28a1 · inbound

Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs cites this paper.

Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:48.072540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:31:17.420639Z digest=sha256:8be7d62ca00e12c86de1ae1ecca8c24c4f001435e98db0fd1fa5791dd40bab42

Observation a5243836-dd44-4be2-b867-29d28b90e009 · inbound

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size cites this paper.

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:46:06.499168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:18:33.458143Z digest=sha256:dcd96d433c4342b91a332ab58bcd386f406e47a70768aac1cc3770ad996c921a

Observation e53f190d-6c90-453b-8425-d49eb776e48f · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:39.819716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:1ac4c34f11fc1d366a4935b5a3c1b6ba0d0e3247f43add87e28d843820c49ae0

Observation 33fd6283-7502-4d89-822b-da9af1e96e67 · inbound

Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models cites this paper.

Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:12:38.900193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:10:52.582492Z digest=sha256:d22c8908dc0248eda622906a7bca4dc6213968608acdc5bbbf1e385d4ded34a9

Observation e9e163fd-e946-4b77-a4fc-472196530f2d · inbound

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions cites this paper.

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:14.808607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:33:05.952601Z digest=sha256:8e36b331b86f251d3d597071d72b478998e7d5004b586eddf16bf11ec8a5f044

Observation fd334a5c-971e-46d0-917f-5e35d581e4cb · inbound

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions cites this paper.

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:43.750585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:43.750585Z digest=sha256:2d88599a274c659567d958e847a96badccccb1ddcc2a8940020efc58f2461867

Observation e4d79cee-fbab-475c-a54c-8e6d81140666 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:36:44.918374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:c2f83e6e9522c2e3b309589223a05b71df136f98eb70c448c8a1fdca009d8f75

Observation 83636f29-a75c-4049-8d5e-065e442f67a7 · inbound

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training cites this paper.

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:46:56.727189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:53:04.715108Z digest=sha256:4c7642fbb8367bab313bf075f020f4439c5c99ec11ab438a3447cbee5ceb9e1c

Observation 345089cb-143d-4052-8b8a-99f84689ca1e · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-26T15:39:33.213970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T15:35:51.654392Z digest=sha256:32038c261a762bca8dbeab6d7492971cbd5af33950db2591dddd615b0630b6b9

Observation 4b39796f-3374-4299-852e-1e98a5359572 · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.386062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T21:51:13.457071Z digest=sha256:7c7f81c2a16d2427da3546f5ddcbb446d2e7034f90db54373f0fe78c3a840fbd

Observation be861fad-ea9a-493a-834d-189e4735834a · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:45.114260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:8786389f2d769a480f42b0f5485d47253d2e1d7c44ba19e6119b912195f4b7ff

Observation 7eccc6d1-5602-4853-ae63-1f8ab990bf7f · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.707947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:a1fe3ccdfd36178c1ef90772eabc74a2f00c9182f19a4be8af9cdda019b2c547

Observation 970f52ba-dd88-4742-ad29-5b29c5482ca2 · inbound

Sentence-Level Contextual Entrainment in Large Language Models cites this paper.

Sentence-Level Contextual Entrainment in Large Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:09:57.234685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:52:53.959215Z digest=sha256:f7673013b35728fbd38fdd7e47814ba193746580eb9b227d555f2177e4d363e3