REVIEW 14 cited by
A Primer on the Inner Workings of Transformer-based Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.
Forward citations
Cited by 14 Pith papers
-
Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs
Weight-adjusted gradients (weight times gradient) identify sparse LLM parameters whose masking induces rapid collapse and improve several efficiency and editing applications.
-
MICE: Minimal Interaction Cross-Encoders for efficient Re-ranking
MICE is a cross-encoder-derived late-interaction ranker that retains most in-domain effectiveness and beats same-size ColBERT by 5-8 nDCG@10 points while cutting latency up to 4x with precomputed document vectors.
-
Deductive Logic in Language Models: Horizontal vs Vertical Reasoning
A 2-layer, single-head attention-only transformer learns to perform multi-step logical deduction through induction-head circuits for rule completion, chaining, and final decision.
-
Cross-Attention is Half Explanation in Speech-to-Text Models
Cross-attention in speech-to-text models correlates with saliency-based explanations (Pearson r roughly 0.49-0.75 in the best aggregations) but explains only a minority of the variance, so it should complement, not re...
-
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
A new benchmark, TableEval, with 3017 tables in five formats, shows LLMs are robust to table representation but perform worse on scientific tables, with the caveat that the domain gap is confounded by task difficulty.
-
REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering
A VQ-AE-based module-scoring method picks steering locations in LLMs, improving truthfulness and knowledge-selection steering over ITI and SPARE baselines.
-
Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models
PME detects memorized personal information in LLMs and edits the feed-forward layer weights so the model outputs a dummy value instead, reducing extraction attack success while preserving general model quality.
-
Different Speech Translation Models Encode and Translate Speaker Gender Differently
Traditional encoder-decoder speech translation models encode speaker gender in hidden states, while newer adapter-based models largely do not; lower gender encoding tracks with masculine-default translation bias.
-
COMPKE: Complex Question Answering under Knowledge Editing
COMPKE is a new benchmark with 11,924 complex questions that tests knowledge editing through one-to-many relations and logical operations, where existing editing methods often fail.
-
Localizing Persona Representations in LLMs
Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.
-
InTraVisTo: Inside Transformer Visualisation Tool
A GUI tool that decodes hidden states into tokens, visualizes information flow via a Sankey diagram, and supports embedding injection for interactive probing of transformer LLMs.
-
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
Claims a two-stage transformer scaling law (exponential then C^{-1/6}) with matching bounds, but the lower bounds are missing, the exponent is inconsistent (-1/7 vs -1/6), and the law is an artifact of hand-set M = Θ(...
-
Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.
-
Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models
A small BERT model trained on only 117 of 517 curated regulatory relationships selected as confident errors reaches 93% balanced accuracy, outperforming a policy that also includes uncertain correct examples.
Discussion (0). Continue with ORCID to comment.