REVIEW 3 cited by
Anchor function: a type of benchmark functions for studying language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Understanding transformer-based language models is becoming increasingly crucial, particularly as they play pivotal roles in advancing towards artificial general intelligence. However, language model research faces significant challenges, especially for academic research groups with constrained resources. These challenges include complex data structures, unknown target functions, high computational costs and memory requirements, and a lack of interpretability in the inference process, etc. Drawing a parallel to the use of simple models in scientific research, we propose the concept of an anchor function. This is a type of benchmark function designed for studying language models in learning tasks that follow an "anchor-key" pattern. By utilizing the concept of an anchor function, we can construct a series of functions to simulate various language tasks. The anchor function plays a role analogous to that of mice in diabetes research, particularly suitable for academic research. We demonstrate the utility of the anchor function with an example, revealing two basic operations by attention structures in language models: shifting tokens and broadcasting one token from one position to many positions. These operations are also commonly observed in large language models. The anchor function framework, therefore, opens up a series of valuable and accessible research questions for further exploration, especially for theoretical study.
Forward citations
Cited by 3 Pith papers
-
Adaptive Preconditioners Trigger Loss Spikes in Adam
Loss spikes in Adam occur when its second-moment memory decays faster than gradients grow, briefly removing the adaptive brake; a single Hessian-vector product along the gradient direction can flag the onset.
-
An Analysis for Reasoning Bias of Language Models with Small Initialization
Initialization scale controls whether a transformer learns compositional reasoning or memorized mappings, because reasoning tokens acquire more differentiated embeddings early in training.
-
Reasoning Bias of Next Token Prediction Training
Training on all tokens (next token prediction) beats training only on answer tokens (critical token prediction) on small-scale reasoning benchmarks, an effect the authors attribute to noise-induced regularization.
Discussion (0). Continue with ORCID to comment.