Pith. sign in

REVIEW 2 cited by

Identifying Linear Relational Concepts in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08968 v2 pith:YHM6N2I3 submitted 2023-11-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords conceptconceptsdirectionslinearrelationalclassifiersfindfinding
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Transformer language models (LMs) have been shown to represent concepts as directions in the latent space of hidden activations. However, for any human-interpretable concept, how can we find its direction in the latent space? We present a technique called linear relational concepts (LRC) for finding concept directions corresponding to human-interpretable concepts by first modeling the relation between subject and object as a linear relational embedding (LRE). We find that inverting the LRE and using earlier object layers results in a powerful technique for finding concept directions that outperforms standard black-box probing classifiers. We evaluate LRCs on their performance as concept classifiers as well as their ability to causally change model output.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RCStat: A Statistical Framework for using Relative Contextualization in Transformers

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RCStat uses pre-softmax attention logits to define a Relative Contextualization score that improves adaptive KV-cache eviction and attention-head selection for attribution on LLaMA models.

  2. Geospatial Mechanistic Interpretability of Large Language Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLM internal representations of placenames show measurable spatial autocorrelation, and sparse autoencoders extract a small number of geospatially interpretable features.

Pith tools