Pith. sign in

Paper Citation Record · LEDGER

Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2404.05405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05405 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:59.545380Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dae8452d-0891-4b6e-8f99-e09d2dee3682 · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:59.545380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:59.545380Z digest=sha256:f93146b5e60ece0ecd47195a1758839cc3d23b94e9fdfab2d80c069ec3b6f253

Observation fac541c2-6b00-412c-af75-dd51c7efab8e · inbound

Benchmarking and Rethinking Knowledge Editing for Large Language Models cites this paper.

Benchmarking and Rethinking Knowledge Editing for Large Language Models Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:43.086949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:43.086949Z digest=sha256:4914f3b31de75d688fe89164e1abcd5274f124f7fc4081ae55eb04763dcd49df

Observation 6f112d02-2bfd-4ede-9044-a566f22c6636 · inbound

Who Reasons in the Large Language Models? cites this paper.

Who Reasons in the Large Language Models? Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:47:24.755626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:47:24.755626Z digest=sha256:c7fe05d796e33154acd6e2e554958ae207b7bc6b569d568dc8f66758ec7fb30b

Observation 230120bb-4c84-4e7c-ad86-ed7509406b8b · inbound

How much do language models memorize? cites this paper.

How much do language models memorize? Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.752224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.752224Z digest=sha256:9fde1f95993b3d87f92b7231c196c04f0df9bb76519f63d817076fe797a3a6d5

Observation 8363ba87-329c-42c9-9ecc-0a23659bea68 · inbound

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer cites this paper.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.571092Z digest=sha256:0d08c8a3cbd4050600e98f08e34c52f471b591c7509c9fcb95aa735ae65b4a77

Observation 5b8cac3c-1757-41fe-92ee-44ad9b2ec53a · inbound

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research cites this paper.

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:16.928046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:16.928046Z digest=sha256:6257b1bb05779b52b6fac9174aeeb11c76e73dcf91e6654e99d3962dbd385fe1

Observation 12765b3c-a70b-4027-b554-92f4ff607306 · inbound

LRM-1B: Towards Large Routing Model cites this paper.

LRM-1B: Towards Large Routing Model Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:40.991429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:40.991429Z digest=sha256:75af3ff2c6ba844909923d5db5b39da7d2e5b927a5b0a47318be2c7ec4d5b5c6

Observation a2d54293-664b-4967-962c-c01c8d13cfd6 · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:49:03.055714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:17b325c395826fa413c1fa013b3061ddf72aeb1246343e9559a82252c4f08d7d

Observation 2f4a7128-fe57-4182-bffd-7867cd67e99d · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:46.790701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:46.790701Z digest=sha256:f9a0253f09df6b63fc83aa8d8958c0415b9e8ce6d2347bcd7bd455643abd8001

Observation 80db9ad1-86fd-48bb-916b-0bdcf700aa62 · inbound

Unifying Learning Dynamics and Generalization in Transformers Scaling Law cites this paper.

Unifying Learning Dynamics and Generalization in Transformers Scaling Law Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T14:02:56.631725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:02:56.631725Z digest=sha256:9b178fe8f2c880649de7c2013463e705388694804b24ad7ca25104f78a94aedf

Observation bb00fe68-5dc0-4c74-ba4e-e5b3d945a46c · inbound

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory cites this paper.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:38:16.515544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T23:37:33.106390Z digest=sha256:8fba033115a66f48fa8a5aebbe79ecbeb87849963f126b5acdc97e4c272a0e20

Observation 3ca16e85-c092-4cea-870a-802d59fa8e1a · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:59.180654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:7481cd2e8109d9f461da3bc38ce2f4cf6c059c446e407949a5d67525fba9f767

Observation c8b17a7a-e22b-4e42-bde3-aab89807ec82 · inbound

Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval cites this paper.

Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:16:09.968352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T16:23:37.393165Z digest=sha256:10e53aed1935a09cc173108020fc611ec7443503e3f2674d7e7de442b5cf3921

Observation be8cea79-b02a-41a6-9ef6-3da7a203bc97 · inbound

The Statistical Cost of Adaptation in Multi-Source Transfer Learning cites this paper.

The Statistical Cost of Adaptation in Multi-Source Transfer Learning Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 149

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:21:20.732302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:19:05.837824Z digest=sha256:cd06a44b5f885f09f7e963d83c3286e36fa5b797c3cf4e4f58dc6e100c0b1c2e

Observation 6c2c10a4-5353-46a6-9997-5f52857e7eb0 · inbound

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds cites this paper.

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:56:21.694097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:55:28.036044Z digest=sha256:8e832ce280e3c68e4fd2fe40fb57c73db4d14cbe2971633262a2908337519568

Observation ab0cf40e-d3e9-400d-8816-7ab6d0e3cd69 · inbound

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data cites this paper.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.645042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:519ed73fe1d73afefa0b066c520ae98fcfb4c00230d3b0154b29a9c209317538

Observation a3c588df-af3e-468f-b981-11cb2ccf4843 · inbound

Geometric Factual Recall in Transformers cites this paper.

Geometric Factual Recall in Transformers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:16.765589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T04:50:17.479689Z digest=sha256:2e51ec085d1c5bb0d9adaebd1c3456eece3d41d03b12160b108fc4afec3acf3f

Observation b2ddf72e-668a-46af-9637-d1fe8700ced0 · inbound

Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency cites this paper.

Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.729220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:59:32.371351Z digest=sha256:887eb707a0173206a45da57ccaf0d4e94d7ee4ee7f3633d7e5587d181bba56dc

Observation 361ac1e7-15c4-419e-9356-0939618aa96a · inbound

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay cites this paper.

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:02.010186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T23:15:45.174086Z digest=sha256:742545cb28275d2b02b1beea118779a5b17ebbd6ef0f40d55339d65dca3eaec0

Observation 30009f3b-5a48-4c83-89a9-bfd29aeec2b5 · inbound

Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence cites this paper.

Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.384173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:03:00.055800Z digest=sha256:d53540f558849dc3cab5f0a1bc2c5316eab41e641faf556d1a474dd45f08e399

Observation 34fcdca2-4fc8-4dec-8db4-b42631fa4d57 · inbound

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy cites this paper.

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.194790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:19:51.893740Z digest=sha256:fe3325064a0f2252a062021e3f097c6c61a87a4c164730d4a489e0e32d23e688

Observation d793f814-5ab8-4bff-9e5c-0c5c149ff787 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:06:44.968032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:5bdd3a5828e115532b32c599babe5a2c1804db2b32436a5e623333fe17757b7f

Observation 208f2e23-d01e-413c-8bc4-d1050aa7904b · inbound

User as Engram: Internalizing Per-User Memory as Local Parametric Edits cites this paper.

User as Engram: Internalizing Per-User Memory as Local Parametric Edits Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:19.521587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T20:37:01.382431Z digest=sha256:4ff80eda1cebb5665e03c31ff5289852eb99599485607488dc3a2957615e5b8d

Observation 5691762a-6d23-4559-9c4a-454c35006e6d · inbound

Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization cites this paper.

Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:21:29.896353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:21:29.896353Z digest=sha256:ccf05ef7492fb8cb0452d4a19e7d33d266a3583aecfbb07e8e2a271b3ec2b8bd

Observation 77a5b3a1-41fd-43c3-a040-0b17199b8e21 · inbound

Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA cites this paper.

Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:29:14.690542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:29:14.690542Z digest=sha256:d6ce05fb346467ddfc44a81ba151c1784463bf73c07e20e8b05e36f200cf3ab9