REVIEW 3 major objections 3 minor
Weight writes can store usable facts from broad data, but later writes make earlier facts unreachable no matter the intervention tested.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 09:02 UTC pith:T5H7FFRB
load-bearing objection Abstract-only: sequential weight writes create usable but question-keyed knowledge that later writes make unreachable; context wins for accumulation, methods still uncheckable. the 3 major comments →
Can a Language Model Learn Facts Continually in Its Weights?
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Broad training data can create usable knowledge inside a model's weights, and a frozen reference model can preserve capability, but after sequences of later weight writes no tested intervention keeps earlier facts reachable; forgotten facts retain most of their log-probability gain yet become question-keyed and redirected by newer writes, so composition and long-term retention of facts require context rather than the weights.
What carries the argument
Sequential weight writes of invented facts into Qwen3, tracked from creation through 20-100 later writes and scored on five held-out question types against a reference that receives the fact only in its prompt; breadth of study data versus bare statements is the variable that determines recitation-to-use gap and later retention.
Load-bearing premise
That invented facts trained into one model family, scored on five kinds of held-out questions, with the original model given the fact in its prompt as reference, are a valid proxy for whether real knowledge is stored, usable, and retained under continual weight writes.
What would settle it
Train a sequence of real, non-invented facts into the same models using broad study data, then measure whether earlier facts remain answerable on the same five held-out question types after twenty later writes; recovery above the reported 46 percent or success of any intervention that restores reachability would contradict the central claim.
If this is right
- Bare-statement continual learning yields near-total loss of earlier facts after only twenty writes, so simple fact dumps cannot accumulate knowledge in weights.
- Diverse restatements create knowledge that is usable and more durable, yet still fails to stay reachable once later writes arrive.
- Behavioural forgetting does not erase stored log-probability; the fact remains but is no longer retrieved by its original questions.
- Unrelated abilities degrade in proportion to KL divergence from the original model, independent of how the earlier fact was stored.
- For any task that requires composing facts or keeping them after later updates, the reliable route is to supply the facts in context rather than rely on weight storage.
Where Pith is reading between the lines
- If the same pattern holds for real knowledge, production systems that continually fine-tune on new facts will need an external memory or retrieval layer rather than pure weight accumulation.
- The observation that wrong answers contain the most recent fact suggests interference is largely a routing or keying problem, inviting tests that re-key older facts without rewriting them.
- A frozen reference that restores capability when the forgotten fact is re-supplied in the prompt points to hybrid designs that keep a clean base model and inject facts only at inference time.
- The recitation-to-use gap reduction without showing conclusions may generalise to other forms of synthetic data diversity beyond the restatements tested here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies whether continual weight updates can accumulate usable knowledge in language models. Invented facts are written into Qwen3 models and tracked through sequences of 20–100 later writes, evaluated on five types of held-out questions against a reference in which the original model receives the fact in its prompt. The abstract reports that training-data breadth determines knowledge quality: bare-statement training yields recitation, while diverse restatements shrink the recitation-to-use gap (27.4 → 5.4 points). After twenty sequential writes, bare facts retain ~1% accuracy versus ~46% for broad-study facts. Forgotten facts retain most of the log-probability added by their write, and under bare training ~70% of wrong answers contain the most recent fact; re-supplying a forgotten fact in context recovers 77–80%. Damage to unrelated abilities tracks KL from the original model. The authors conclude that knowledge is stored but question-keyed, that no tested weight-based intervention keeps earlier facts reachable, and that context is the reliable channel when facts must be composed or survive later writes.
Significance. If the experimental design holds under full scrutiny, the work is a substantial empirical contribution to continual learning and knowledge editing. It cleanly separates recitation from use, documents behavioural forgetting without erasure, quantifies interference under sequential writes, and supplies a strong negative result on weight-based retention interventions. The breadth-of-data finding and the frozen-reference preservation observation are practically useful. The claim that context, not weights, is the reliable channel for composition and long-horizon retention would reframe research priorities in the area. Credit is due for the multi-type held-out suite, the prompt-reference baseline, and the explicit tracking of log-probability retention alongside accuracy.
major comments (3)
- The central negative claim—that no tested weight intervention keeps earlier facts reachable—rests on the validity of invented facts, the five held-out question types as probes of use/composition rather than recitation, and the fairness of the original-model-with-fact-in-prompt reference. The abstract alone cannot establish that surface-form confounds are controlled, that the five types isolate composition, or that the reference fairly separates weight storage from context. These design choices are load-bearing: if any fails, the reported retention numbers (1% vs 46%, 70% recent-fact contamination, 77–80% context recovery) no longer support the conclusion. Full methods, item construction, and ablation of the question suite are required before the claim can be accepted.
- The abstract asserts that later writes cause interference 'regardless of how the earlier fact was stored' and that damage tracks KL divergence from the original model. Without the full experimental matrix (which interventions, which KL regimes, statistical uncertainty, and whether held-out items truly isolate use), it is impossible to judge whether the interference result generalises or is an artefact of the particular write schedule and model scale. This is load-bearing for the claim that weight writes cannot support accumulation.
- The strong conclusion that 'the reliable channel is context rather than the weights' is an existence claim over the interventions tested. The abstract does not enumerate those interventions or their local-measurement accuracy. A complete list, with negative results stated per method, is needed so readers can assess whether the search was exhaustive enough to support the generalisation.
minor comments (3)
- Abstract is carefully written and reports specific percentages and design elements; once the full paper is available, ensure that every numeric claim in the abstract is traceable to a table or figure with error bars or confidence intervals.
- The interpretive construct 'question-keyed knowledge' should be defined operationally in the main text (e.g., via the five question types and the log-probability retention result) so it is not read as a free theoretical entity.
- Clarify data-exclusion rules, hyperparameter choices, and whether any facts or questions were filtered post hoc; these are standard reproducibility items that the abstract cannot supply.
Circularity Check
No significant circularity: empirical continual-learning measurements against held-out questions and an external prompt reference, with no derivation that reduces by construction to its inputs.
full rationale
This abstract-only paper reports experimental measurements of invented facts written into Qwen3 weights, scored on five types of held-out questions, with the original model given the fact in its prompt as the reference baseline. There are no equations, fitted parameters renamed as predictions, uniqueness theorems, or self-citation chains that force the central claims by construction. Retention numbers (e.g., 1% bare vs 46% broad after twenty writes; 77–80% recovery when a forgotten fact is re-supplied in context) are outcomes of sequential training and evaluation, not tautological restatements of fitted targets. Defining “knowledge” via the same held-out question suite used for scoring is ordinary experimental self-consistency, not equation-level circularity. Proxy-validity concerns about invented facts or question design are correctness risks outside the circularity criteria. With only the abstract available and no load-bearing self-definitional or self-citation reduction quotable from the text, the score is 0 and steps is empty.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Invented facts written into model weights are a valid proxy for real-world factual knowledge acquisition and retention.
- domain assumption Accuracy on five types of held-out questions distinguishes usable knowledge from mere recitation.
- domain assumption The original model given the fact in its prompt is an appropriate reference for what successful knowledge should look like.
- domain assumption Sequential fine-tuning writes of later facts are a fair operationalization of continual weight-based learning.
invented entities (1)
-
question-keyed knowledge (interpretive construct)
independent evidence
read the original abstract
Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided. We follow invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later writes, using held-out questions of five types, with the original model given the fact in its prompt as the reference. Across these experiments, the breadth of the training data determines the kind of knowledge created. Bare-statement training produces recitation, while diverse restatements reduce the recitation-to-use gap from 27.4 to 5.4 points without showing the model a conclusion. This difference carries into later writes: after twenty sequential writes, bare-statement facts retain 1% accuracy while facts written from broad study data retain 46%. We also find that facts can be behaviourally forgotten without being erased. Forgotten facts keep most of the log-probability added by their write, and under bare-statement training 70% of wrong answers about them contain the most recently written fact. The same writes barely degrade the model's use of facts in context, and a forgotten study fact supplied in the prompt recovers to 77-80% on its questions. These results describe knowledge that is stored but question-keyed: later writes redirect the questions that reached it. Damage to unrelated abilities tracks KL divergence from the original model, and the later writes cause interference regardless of how the earlier fact was stored. Broad data can create usable knowledge, and a frozen reference can preserve capability, but no intervention we tested, including those built on accurate local measurements of each write, keeps earlier facts reachable. When facts must be composed or survive later writes, the reliable channel is context rather than the weights.
Figures
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.