Pith. sign in

Paper Citation Record · LEDGER

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2508.10975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10975 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:24:31.154102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T16:54:58.206077Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0433471-6744-4cde-8e03-160a390f6dfa · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:58.134824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:58.134824Z digest=sha256:9a13195901ac71533c26df41b3ddd329f8793f8b220c3cc4e56d5de26af88ac0

Observation 62560294-e859-4955-9fd9-d16622566338 · inbound

Action-guided generation of 3D functionality segmentation data cites this paper.

Action-guided generation of 3D functionality segmentation data BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:51:31.948881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T04:49:06.269256Z digest=sha256:a810af521393a2dee1444353bd1ce79e7df3179bac46a0442a07a2b365e2da12

Observation 4e028ef2-af96-4ef3-8b1b-581fe2349e5c · inbound

Linguistics and Human Brain: A Perspective of Computational Neuroscience cites this paper.

Linguistics and Human Brain: A Perspective of Computational Neuroscience BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 197

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:46.300889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:46.300889Z digest=sha256:955c965bc4db70e5181959046df51036191dfc581627434e4b8c0e277c910d14

Observation 0f7db37f-fd7f-4ca6-ab9c-c886dea62ed6 · inbound

Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale cites this paper.

Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:15.130856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T11:44:58.613981Z digest=sha256:6957fad2c0cf89a125bac5d01038ebc2d824a1ad00a937b83e037d6d0e48df90

Observation b032f9a5-3ba8-490f-998e-2fa7cfccb265 · inbound

Understanding Data Temporality Impact on Large Language Models Pre-training cites this paper.

Understanding Data Temporality Impact on Large Language Models Pre-training BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T16:54:58.207894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T16:53:01.583415Z digest=sha256:ea2b7feb4bde2ee4f8ccb8184ebea2bf18470aad367847832aef90abbaf47a93

Observation 4f668e0a-0926-489d-8397-14e58271fe9e · inbound

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention cites this paper.

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.203279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T08:39:14.327838Z digest=sha256:237faa346752c9bfd3ea22d9e075e7c200f3ba20a798bd7f9eb540edb72e9fa8

Observation 04a89781-b13e-4781-bf90-88f658abbf78 · inbound

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data cites this paper.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T07:01:45.805431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:01:45.805431Z digest=sha256:e3c70c59dcd0a0a6106ce38abb57016241443c163cc5a68110aceb2d009d7dd7

Observation afc736d1-078d-4fbd-a7b6-8bed5d1cf7f6 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.696790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.696790Z digest=sha256:54ffa65b2e5efa4bd4a45546467e7d41ef141b562371f98b35360dadf11a669b

Observation 582fe0a1-72d9-4dce-bb84-2ee3e8bc055d · inbound

Deep Research Pretraining via Predictive Navigation cites this paper.

Deep Research Pretraining via Predictive Navigation BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:24:31.154102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:24:31.154102Z digest=sha256:1e529d55ec550551fee9bd3f2d0f9b4a773f6ca240f8c464dc316e00d994999d

Observation 1e9d02fe-702a-4a6d-9bd9-36f582b86f14 · inbound

The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora cites this paper.

The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:46:08.585204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:46:08.585204Z digest=sha256:e9ab9efb99e8f9d8bda0368e837973bd5583e93749da984a114ec1af9cb8649e