Pith. sign in

Paper Citation Record · LEDGER

Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2304.03208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.03208 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.474327Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

24
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 72709979-5cf3-409f-93c8-87bf672b1828 · inbound

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only cites this paper.

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:43:45.805619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:43:45.770157Z digest=sha256:6b41a11fcc1fbb3fc06b1d4439ab755df02bfacd906f549c3461a602ec6a9b05

Observation 11ce8e20-346e-4ee5-9948-e01920c03746 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 273

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.073190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:278401340aac93e4801796f1c305cc071c82a80d1a2f2e575e34e6efde3957b4

Observation 7ded08bb-9350-4ace-a73c-fa0057af11eb · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:00:53.450722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:0ba49be00b6119baf1403c0084ef89ddd1bc327bb96b2310a67da129aa8a9d10

Observation 7f2cbd0f-da97-422e-8f01-f1548ff192de · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.423436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:415a89bc9a76e896493c6bdaa5326235c870ecdbd15f0634e959bb828f733074

Observation 8d69a650-a8bd-4bfe-bada-f1598e01d261 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.474327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.474327Z digest=sha256:09d833a0ad3fe689e4a8323af37039953d730379f5f89f59be25561ecf9af9d1

Observation 0d10b42b-515c-40a5-b47a-0fafba6430f3 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:23.116345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:23.116345Z digest=sha256:87275c964d4d9f424f9431d5244e3385c5cd18644d6e08a560ff21301617ff91

Observation df1ccc47-c817-46e5-b64c-4ed0e368cc4e · inbound

MuLoCo: Muon is a practical inner optimizer for DiLoCo cites this paper.

MuLoCo: Muon is a practical inner optimizer for DiLoCo Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:26.558466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:26.558466Z digest=sha256:d12e49f9daeab3b4b29aa95f446413fa2674b4f658c5de5ad71314035386707b

Observation 7b19a67a-b969-4e9b-bd05-f3e374efaaa7 · inbound

Basis Transformers for Multi-Task Tabular Regression cites this paper.

Basis Transformers for Multi-Task Tabular Regression Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:18.794719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:18.794719Z digest=sha256:6c44715d49b6f2204c97ed2a38484155dd254f2a6e6e784b478eb95baec2e3ba

Observation 4c52a78a-a6e0-4ad8-a59b-a3e36ad36ad0 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.928577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.928577Z digest=sha256:f1f8a8abc1f2c3bae01b478d76f206103a821c35fe535e6800f69e5f0ab60b71

Observation 696b93b9-96f5-49dc-af04-c4a478f741f8 · inbound

Similarity Field Theory: A Mathematical Framework for Intelligence cites this paper.

Similarity Field Theory: A Mathematical Framework for Intelligence Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:41:30.649201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:36:49.543497Z digest=sha256:a958bab6e7ab9edf6caf7ed7242aaca3849910ec420f4ddbb7a2fd18fcd8f811

Observation f6d2d069-c6f6-451d-9f9d-23f49452f5d0 · inbound

SpaDA: A Spatial Dataflow Architecture Programming Language cites this paper.

SpaDA: A Spatial Dataflow Architecture Programming Language Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:20:23.248975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:18:38.464461Z digest=sha256:6d7d6abcf8fb851256724fb2c65f9f45fec26ebd45476795ba840036e5822a9c

Observation 1c9812f9-a439-4926-ad6d-d58639597ea3 · inbound

Spectral Condition for $\mu$P under Width-Depth Scaling cites this paper.

Spectral Condition for $\mu$P under Width-Depth Scaling Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:06:25.514696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T18:03:31.202134Z digest=sha256:7fc72df806720106f6e7d5ca3d372297b3a7421cf91510e161c51020a744cfd2

Observation 3030597b-90d3-49d4-b542-8bcf4d65b7e0 · inbound

Learning the Signature of Memorization in Autoregressive Language Models cites this paper.

Learning the Signature of Memorization in Autoregressive Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:53:11.472570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:53:10.396785Z digest=sha256:967982415ea852546c3179980ced09d5485f3ed17180ba8aec1823b4ac3fccad

Observation 397fe58b-91bd-4d8c-832d-adabb3ac28a1 · inbound

Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs cites this paper.

Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:48.072540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:31:17.420639Z digest=sha256:9372fac02e6df2e5a5f86489bf72861cfa15b38c2819c75d88a801bef022718a

Observation a5243836-dd44-4be2-b867-29d28b90e009 · inbound

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size cites this paper.

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:46:06.499168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:18:33.458143Z digest=sha256:e5cdd481a11152660e6dc25210311c7f7ecd842fa78b854e3a675bb6e70c7032

Observation e53f190d-6c90-453b-8425-d49eb776e48f · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:39.819716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:5b4892f2fc9e6b03cefb78fcfbfbad4d4125c6c0159f3774c43dad7ba674a4e4

Observation 33fd6283-7502-4d89-822b-da9af1e96e67 · inbound

Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models cites this paper.

Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:12:38.900193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T16:10:52.582492Z digest=sha256:a90311fbf2aa5d3d46068dde01c41ca188ef3ceb84c0e5b2c514804be259bc73

Observation e9e163fd-e946-4b77-a4fc-472196530f2d · inbound

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions cites this paper.

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:14.808607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:33:05.952601Z digest=sha256:bfd661f29e5a321f8d89c167e380bc4736212cdebc2663de59bcb7b6c8963c97

Observation fd334a5c-971e-46d0-917f-5e35d581e4cb · inbound

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions cites this paper.

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:43.750585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:43.750585Z digest=sha256:2d88599a274c659567d958e847a96badccccb1ddcc2a8940020efc58f2461867

Observation e4d79cee-fbab-475c-a54c-8e6d81140666 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:36:44.918374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:81d32c92f03e71e13529dc9038c18eb8ed69ded8a71f04e252caf14513652026

Observation 83636f29-a75c-4049-8d5e-065e442f67a7 · inbound

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training cites this paper.

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:46:56.727189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:53:04.715108Z digest=sha256:8c7aa4f585d82b0ee1a87d3729815ae94522ee37754a676af1a086ac1c60dbd7

Observation 345089cb-143d-4052-8b8a-99f84689ca1e · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-26T15:39:33.213970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T15:35:51.654392Z digest=sha256:0cefa47e1c0ca61b779883924c266e5c52be4382c27f61ac3591fcc552c00a3a

Observation 4b39796f-3374-4299-852e-1e98a5359572 · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.386062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T21:51:13.457071Z digest=sha256:e2624c5353b69a1b5401643b2a3cacac57e90d3c18d90626f99a9da62ba69a50

Observation be861fad-ea9a-493a-834d-189e4735834a · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:45.114260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:bb908c5f1281908df41387cdfb7c936a859ea8915bec06bb9a585b6793d7ba5b

Observation 7eccc6d1-5602-4853-ae63-1f8ab990bf7f · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.707947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:33c5229075ef5eefae0d91f9946b1132b890d1df4d0ad7884dbb87ca5fa54124

Observation 970f52ba-dd88-4742-ad29-5b29c5482ca2 · inbound

Sentence-Level Contextual Entrainment in Large Language Models cites this paper.

Sentence-Level Contextual Entrainment in Large Language Models Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:09:57.234685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T00:52:53.959215Z digest=sha256:f74421bf52731ca4dabd0f0fb199500a764ff9c3a4260c48a59ad2b2e9abee97