Pith. sign in

REVIEW 14 cited by

A Primer on the Inner Workings of Transformer-based Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00208 v3 pith:5L7B7RSI submitted 2024-04-30 cs.CL

classification cs.CL
keywords modelsinnerlanguageworkingsareaprimerresearchtransformer-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area. This primer provides a concise technical introduction to the current techniques used to interpret the inner workings of Transformer-based language models, focusing on the generative decoder-only architecture. We conclude by presenting a comprehensive overview of the known internal mechanisms implemented by these models, uncovering connections across popular approaches and active research directions in this area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Weight-adjusted gradients (weight times gradient) identify sparse LLM parameters whose masking induces rapid collapse and improve several efficiency and editing applications.

  2. MICE: Minimal Interaction Cross-Encoders for efficient Re-ranking

    cs.IR 2026-02 conditional novelty 6.0 of 10

    MICE is a cross-encoder-derived late-interaction ranker that retains most in-domain effectiveness and beats same-size ColBERT by 5-8 nDCG@10 points while cutting latency up to 4x with precomputed document vectors.

  3. Deductive Logic in Language Models: Horizontal vs Vertical Reasoning

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A 2-layer, single-head attention-only transformer learns to perform multi-step logical deduction through induction-head circuits for rule completion, chaining, and final decision.

  4. Cross-Attention is Half Explanation in Speech-to-Text Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Cross-attention in speech-to-text models correlates with saliency-based explanations (Pearson r roughly 0.49-0.75 in the best aggregations) but explains only a minority of the variance, so it should complement, not re...

  5. Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark, TableEval, with 3017 tables in five formats, shows LLMs are robust to table representation but perform worse on scientific tables, with the caveat that the domain gap is confounded by task difficulty.

  6. REAL: Reading Out Transformer Activations for Precise Localization in Language Model Steering

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A VQ-AE-based module-scoring method picks steering locations in LLMs, improving truthfulness and knowledge-selection steering over ITI and SPARE baselines.

  7. Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models

    cs.CR 2025-06 conditional novelty 6.0 of 10

    PME detects memorized personal information in LLMs and edits the feed-forward layer weights so the model outputs a dummy value instead, reducing extraction attack success while preserving general model quality.

  8. Different Speech Translation Models Encode and Translate Speaker Gender Differently

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Traditional encoder-decoder speech translation models encode speaker gender in hidden states, while newer adapter-based models largely do not; lower gender encoding tracks with masculine-default translation bias.

  9. COMPKE: Complex Question Answering under Knowledge Editing

    cs.CL 2025-06 conditional novelty 6.0 of 10

    COMPKE is a new benchmark with 11,924 complex questions that tests knowledge editing through one-to-many relations and logical operations, where existing editing methods often fail.

  10. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  11. InTraVisTo: Inside Transformer Visualisation Tool

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A GUI tool that decodes hidden states into tokens, visualizes information flow via a Sankey diagram, and supports embedding injection for interactive probing of transformer LLMs.

  12. Unifying Learning Dynamics and Generalization in Transformers Scaling Law

    cs.LG 2025-12 reject novelty 4.0 of 10

    Claims a two-stage transformer scaling law (exponential then C^{-1/6}) with matching bounds, but the lower bounds are missing, the exponent is inconsistent (-1/7 vs -1/6), and the law is an artifact of hand-set M = Θ(...

  13. Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations

    cs.CL 2025-10 conditional novelty 4.0 of 10

    LLM explanations split into local and mechanistic tracks; the paper argues they are trustworthy only if they pass causal and contrastive stress tests, adapt to the explainee, and satisfy eight trust principles.

  14. Reconstructing Biological Pathways by Applying Selective Incremental Learning to (Very) Small Language Models

    q-bio.MN 2025-07 conditional novelty 4.0 of 10

    A small BERT model trained on only 117 of 517 curated regulatory relationships selected as confident errors reaches 93% balanced accuracy, outperforming a policy that also includes uncertain correct examples.

Pith tools