Pith. sign in

REVIEW 2 cited by

Does Transformer Interpretability Transfer to RNNs?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05971 v1 pith:VANKLPOK submitted 2024-04-09 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords rnnsarchitectureselicitinginterpretabilitylanguagelatentmodelsoutputs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of language modeling perplexity and downstream evaluations, suggesting that future systems may be built on completely new architectures. In this paper, we examine if selected interpretability methods originally designed for transformer language models will transfer to these up-and-coming recurrent architectures. Specifically, we focus on steering model outputs via contrastive activation addition, on eliciting latent predictions via the tuned lens, and eliciting latent knowledge from models fine-tuned to produce false outputs under certain conditions. Our results show that most of these techniques are effective when applied to RNNs, and we show that it is possible to improve some of them by taking advantage of RNNs' compressed state.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Understanding What State Space Models Learn About Code

    cs.AI 2026-02 conditional novelty 6.0 of 10

    SSM code models capture code syntax and semantics better than Transformers before fine-tuning, forget short-range structure when fine-tuned on type inference, and an added high-frequency path or more kernels recovers ...

  2. Reproducing Recurrent Transformers: The CoTFormer

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A reproduction of CoTFormer confirms its perplexity results, finds its adaptive-compute claims fragile, and shows looped computation benefits p-hop retrieval but not inductive counting.

Pith tools