Pith. sign in

Paper Citation Record · LEDGER

Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2404.05405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05405 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:25:25.286028Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8807e8bb-2dfb-425a-890a-c20c02444a30 · inbound

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers cites this paper.

Paying Attention to Facts: Quantifying the Knowledge Capacity of Attention Layers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:25.286028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:25:25.286028Z digest=sha256:957432e60cdb1ab6ff9f597c7eb678da12281393f24248b510b52861419de1fe

Observation 53abca1b-6c74-4c3c-ba55-f33eee8ce64b · inbound

Do we really have to filter out random noise in pre-training data for language models? cites this paper.

Do we really have to filter out random noise in pre-training data for language models? Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.637076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.637076Z digest=sha256:4e8db06288e38134adc22a5d3d3ddd6607882bd1c01330abd994a107eec31b90

Observation dae8452d-0891-4b6e-8f99-e09d2dee3682 · inbound

Enhancing LLMs via High-Knowledge Data Selection cites this paper.

Enhancing LLMs via High-Knowledge Data Selection Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:59.545380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:59.545380Z digest=sha256:7b2f5e952ce6a01d68313bf574f756c84c6f5ebcd899ac8549ab7a5a8049b370

Observation fac541c2-6b00-412c-af75-dd51c7efab8e · inbound

Benchmarking and Rethinking Knowledge Editing for Large Language Models cites this paper.

Benchmarking and Rethinking Knowledge Editing for Large Language Models Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:43.086949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:43.086949Z digest=sha256:730eab1265e9ee9f987c75bcf69788e5d15c68ffbb35039278900f4f0a754b71

Observation 6f112d02-2bfd-4ede-9044-a566f22c6636 · inbound

Who Reasons in the Large Language Models? cites this paper.

Who Reasons in the Large Language Models? Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:47:24.755626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:47:24.755626Z digest=sha256:5d61b151db662dc2fd580df40c0fa9f28a80c323574608194e23fe9bf749e4a4

Observation 230120bb-4c84-4e7c-ad86-ed7509406b8b · inbound

How much do language models memorize? cites this paper.

How much do language models memorize? Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:35.752224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:35.752224Z digest=sha256:ca7fff37d7644f33db087eb0cc04e46981258859360407219482c0d3f5d96831

Observation 8363ba87-329c-42c9-9ecc-0a23659bea68 · inbound

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer cites this paper.

Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:22.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:57:22.571092Z digest=sha256:8eadeb10d2e8df0a744daca4ab73d709bfed12192f4bc9079022b136b5eed7a8

Observation 5b8cac3c-1757-41fe-92ee-44ad9b2ec53a · inbound

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research cites this paper.

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:16.928046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:16.928046Z digest=sha256:1146d19b79f3acb144abe63a609b23e990ede813b556ed7bc9649109f363d193

Observation 12765b3c-a70b-4027-b554-92f4ff607306 · inbound

LRM-1B: Towards Large Routing Model cites this paper.

LRM-1B: Towards Large Routing Model Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:20:40.991429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:20:40.991429Z digest=sha256:af6f04680fe7ec6ffc5e9947757507506b284f53bd672b9714501c024e58cd0d

Observation a2d54293-664b-4967-962c-c01c8d13cfd6 · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:49:03.055714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:66278a1a918cd3af9822ba9e54ab22324d8e74eb7883d975aa694f813263fccc

Observation 2f4a7128-fe57-4182-bffd-7867cd67e99d · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:46.790701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:46.790701Z digest=sha256:f9a0253f09df6b63fc83aa8d8958c0415b9e8ce6d2347bcd7bd455643abd8001

Observation 80db9ad1-86fd-48bb-916b-0bdcf700aa62 · inbound

Unifying Learning Dynamics and Generalization in Transformers Scaling Law cites this paper.

Unifying Learning Dynamics and Generalization in Transformers Scaling Law Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T14:02:56.631725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:02:56.631725Z digest=sha256:2ffa4997f561486e5663047283ddcdac0a6031cb895384f7592a0570a1157954

Observation bb00fe68-5dc0-4c74-ba4e-e5b3d945a46c · inbound

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory cites this paper.

Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:38:16.515544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T23:37:33.106390Z digest=sha256:20b26170799cf310a915d8cedf3b0fcfafc59cf8bc6b028a037cfca8b41638d2

Observation 3ca16e85-c092-4cea-870a-802d59fa8e1a · inbound

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts cites this paper.

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:59.180654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:42:31.465077Z digest=sha256:a4645a684c95237d506ac753910c19380bbf1556dede2682b92682d02bbdac8d

Observation c8b17a7a-e22b-4e42-bde3-aab89807ec82 · inbound

Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval cites this paper.

Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:16:09.968352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T16:23:37.393165Z digest=sha256:693faca00f350123bfcbdfc718ad02d0d048c778f768e7e66ef0ef161fe51a35

Observation be8cea79-b02a-41a6-9ef6-3da7a203bc97 · inbound

The Statistical Cost of Adaptation in Multi-Source Transfer Learning cites this paper.

The Statistical Cost of Adaptation in Multi-Source Transfer Learning Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 149

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:21:20.732302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T04:19:05.837824Z digest=sha256:bc52884b005f8239e4a264207631cd4948e0ba35384c9d9fb41b5f3c90786989

Observation 6c2c10a4-5353-46a6-9997-5f52857e7eb0 · inbound

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds cites this paper.

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:56:21.694097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:55:28.036044Z digest=sha256:33aebe9ce8c6ee7b058efd64b3ce521357b027e344e047d6572e99c3bdae0146

Observation ab0cf40e-d3e9-400d-8816-7ab6d0e3cd69 · inbound

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data cites this paper.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.645042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:db221a1b5f5b5f0f9dae15f5660932a9c5bb89864c1f8be173f04b8810711f12

Observation a3c588df-af3e-468f-b981-11cb2ccf4843 · inbound

Geometric Factual Recall in Transformers cites this paper.

Geometric Factual Recall in Transformers Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:16.765589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T04:50:17.479689Z digest=sha256:5542aa9869482a229353212eecbb291a5506bf5032f8bf6fa3216b1261ec301e

Observation b2ddf72e-668a-46af-9637-d1fe8700ced0 · inbound

Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency cites this paper.

Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:03:13.729220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:59:32.371351Z digest=sha256:12af51aebb8b29ee01eff68de8235115b08ca3fae9d4eb502c251383a62e403b

Observation 361ac1e7-15c4-419e-9356-0939618aa96a · inbound

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay cites this paper.

Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:02.010186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:15:45.174086Z digest=sha256:2dd669092d24d13a13676647a75409bde0a278f43e7ea66b95a3ddc08b0d4296

Observation 30009f3b-5a48-4c83-89a9-bfd29aeec2b5 · inbound

Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence cites this paper.

Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.384173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T19:03:00.055800Z digest=sha256:8c993eab6a171ab6cb52d09ccefbaafb248124e264569f7142a07dc56c3b2e09

Observation 34fcdca2-4fc8-4dec-8db4-b42631fa4d57 · inbound

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy cites this paper.

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.194790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:19:51.893740Z digest=sha256:a2703a406ed61715cbe420bf83676d846a98b449427d2688cb4bbe52498ab10b

Observation d793f814-5ab8-4bff-9e5c-0c5c149ff787 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:06:44.968032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:f0eae5e0480aa6c38925107351e4b656ca6b1ccb8873e5f5c54048896b326c4d

Observation 208f2e23-d01e-413c-8bc4-d1050aa7904b · inbound

User as Engram: Internalizing Per-User Memory as Local Parametric Edits cites this paper.

User as Engram: Internalizing Per-User Memory as Local Parametric Edits Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:19.521587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T20:37:01.382431Z digest=sha256:2eaa7004514e5db5933d10af0e7dd1464923686ca34c92c8b40d5a887c950cac

Observation 5691762a-6d23-4559-9c4a-454c35006e6d · inbound

Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization cites this paper.

Local Redundancy: An Information-Theoretic Measure of Plasticity from Synthetic Memorization Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:21:29.896353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:21:29.896353Z digest=sha256:37936bba68d59b24c1a025d70062fc30f2fa657c4d9942eb736308c8d9ec77c9

Observation 77a5b3a1-41fd-43c3-a040-0b17199b8e21 · inbound

Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA cites this paper.

Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:29:14.690542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:29:14.690542Z digest=sha256:d6ce05fb346467ddfc44a81ba151c1784463bf73c07e20e8b05e36f200cf3ab9