Pith. sign in

Paper Citation Record · LEDGER

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2506.05166.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05166 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:45.002961Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:04:06.753097Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T21:18:00.318407Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3772aed2-6acb-4a08-a69c-7a524d6ca671 · outbound

This paper cites online" 'onlinestring :=.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.768165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.768165Z digest=sha256:a2fecdf2fcd7872eed6322ee8d1dd2e2dff5bdb6370959bf07cf0e69a5ee0028

Observation e06ece96-14c5-464e-834d-dc7ad7b63282 · outbound

This paper cites write newline.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.778131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.778131Z digest=sha256:ad95178e372fda9d43283222bf93a93d63dd4cb9eaf6b78d7128de4fe4be7608

Observation 39a7a4ba-b128-4fe7-9ec3-41d8d5ca2ab9 · outbound

This paper cites write newline.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective write newline

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.786484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.786484Z digest=sha256:740a98f7aa5b5e36ff4f6e7d3e8f8a6bc2bbba4a8ab9e12297362c627b466dfc

Observation fcfed349-23e6-4ade-af55-189c39c41452 · outbound

This paper cites Stubbersfield.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Stubbersfield

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.577242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.794108Z digest=sha256:4d0d3769c862ab5b1a36315c3a829a3c7015347c30d9360c3a0b2bb79f125a68

Observation ec4dd6b9-8fdc-4074-b803-81dd46760915 · outbound

This paper cites Science in the age of large language models.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Science in the age of large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.561584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.802060Z digest=sha256:9305aab4e54adbc3e0215c7d700f9466c8659aec94b6b6a1282aeae207db6835

Observation a5cd9e91-6d6b-45aa-9955-547dceb2b4b0 · outbound

This paper cites Quantifying and Reducing Stereotypes in Word Embeddings.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Quantifying and Reducing Stereotypes in Word Embeddings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.807248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.807248Z digest=sha256:096e68153fc98eaf6924760e9795a45721e8b80c380382e807f678393b3b1f25

Observation efc5a07e-07a5-4a00-aa93-730a8a3a21fd · outbound

This paper cites Man is to computer programmer as woman is to homemaker? debiasing word embeddings.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Man is to computer programmer as woman is to homemaker? debiasing word embeddings

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.546048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.814121Z digest=sha256:4e3413f93e8e858121617d7606cb3bdce87eeb10538e7a1e698744fb1bf02fe0

Observation a4c7aba8-cf3a-4836-ba85-db936d2ed7c8 · outbound

This paper cites Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.819483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.819483Z digest=sha256:8257dbeeb0572eecea23054e45a53ac9370d4f6c712602015fd06dc49193b59f

Observation 0b4416e2-4e1f-4970-a8e2-a74e81159d52 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Towards automated circuit discovery for mechanistic interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.827250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.827250Z digest=sha256:d54505fb57d23a29f647a732ba6d5c5dbd3a0bb2c3c3379b6c94a7dc0bf60d7b

Observation 66ee3a01-cf46-4cfe-ae0f-d8639c625ca7 · outbound

This paper cites Gallegos, Ryan A.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Gallegos, Ryan A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.520197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.834649Z digest=sha256:07876b9cb1261ee272fbd236efed00f5fc75ef367f7e5298125db74a0cc60e31

Observation 358dd5c8-2a99-49ff-879c-add0019436fa · outbound

This paper cites Causal abstractions of neural networks.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Causal abstractions of neural networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.841155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.841155Z digest=sha256:c245c316aae1d1567d255e11c5516ce623b03fd1e6bd03eb7f9008c8ce6892d1

Observation 9e3a3c02-6824-4351-b7b7-f80224012029 · outbound

This paper cites Multimodal neurons in artificial neural networks.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Multimodal neurons in artificial neural networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.494054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.849885Z digest=sha256:38bd849c57342868fc5892cb7ea96fc5bbf5bdbdcfc1bd3da20c0839e4426ef3

Observation 9b87056f-b8f7-4591-ae43-061283bd12ea · outbound

This paper cites Localizing Model Behavior with Path Patching.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Localizing Model Behavior with Path Patching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.855872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.855872Z digest=sha256:d9f5e631346de711c899cbd2df7a184c4b2ea71d842b95db1e6df07ff566c2c5

Observation 97959f3a-54bf-4423-9ae8-50f52bc579d7 · outbound

This paper cites C hat GPT based data augmentation for improved parameter-efficient debiasing of LLM s.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective C hat GPT based data augmentation for improved parameter-efficient debiasing of LLM s

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.477923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.861398Z digest=sha256:0149ccb1484d4e3a6a5c5779a339ea6d2f6aa4bc9e6857b85a0636e4f4eb0749

Observation 78ad67f0-2229-4412-89cb-86380b5dfe33 · outbound

This paper cites distilbert-base-uncased-finetuned-sst-2-english (revision bfdd146), 2022.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective distilbert-base-uncased-finetuned-sst-2-english (revision bfdd146), 2022

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.461355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.867114Z digest=sha256:2f442f7a8afbabcd1ad3869ebd91625850c3cc018837d8ab6f45d23beabad4d6

Observation 16c76e03-68e9-4de1-9c26-09f0ed430230 · outbound

This paper cites Shovon, and Gene Kim.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Shovon, and Gene Kim

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.444004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.872004Z digest=sha256:c31eaf988b9a7a3894693bd786de03000627421878185fc95b0fb4d2ac5022df

Observation c8a8dcd2-ad11-46bf-af70-bcb691cfda39 · outbound

This paper cites The impact of debiasing on the performance of language models in downstream tasks is underestimated.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective The impact of debiasing on the performance of language models in downstream tasks is underestimated

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.428479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.877309Z digest=sha256:d8fe9b26febd9ce2d2dbcbae67ebe208b03f180fcd9fbe80c8473c6fc8ca86dc

Observation d0f81c75-6dd2-4d9a-9ad1-a456da917fd3 · outbound

This paper cites Backward Lens: Projecting Language Model Gradients into the Vocabulary Space.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Backward Lens: Projecting Language Model Gradients into the Vocabulary Space

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.882463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.882463Z digest=sha256:548dfc04c1594d433af454664aa9af2ea8dcb9fe0bd46bba3d74a21862124164

Observation f1d12fa1-2227-4053-9c63-9e7738c4a344 · outbound

This paper cites Linear Representations of Political Perspective Emerge in Large Language Models.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Linear Representations of Political Perspective Emerge in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.887656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.887656Z digest=sha256:6e1796cc2b5d2fc9dfed5cec661a06ed93fe62d3308bbe3b763c9c87bc561b74

Observation 02b82d85-c601-4faf-8e6c-89115c2f9854 · outbound

This paper cites Gender bias and stereotypes in large language models.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Gender bias and stereotypes in large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.411906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.893433Z digest=sha256:328785ed2d5093f73f8b9558dc8a0366537ab565d1a5ef25e347e5fd3bc9d070

Observation 791accd6-099f-44b4-a176-3cb679e6cb5b · outbound

This paper cites Sparse feature circuits: Discovering and editing interpretable causal graphs in language models.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Sparse feature circuits: Discovering and editing interpretable causal graphs in language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.395503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.899079Z digest=sha256:40268871b21d7a4674ebfe54d207cf063807eb8d30459db46aeeb53d75f745f7

Observation bf25f2f4-f439-4c5c-bcda-4c18322f57d3 · outbound

This paper cites Locating and editing factual associations in gpt.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Locating and editing factual associations in gpt

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.904438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.904438Z digest=sha256:8923b2c487eb04e338191caa6df9430854e4346e82aa26b0b88cfe9b1e053a6d

Observation 66267b2e-0673-442c-8df7-403023597f77 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Progress measures for grokking via mechanistic interpretability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.911625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.911625Z digest=sha256:1e6d790816e2ac61bee0adc10cbcda8feda659e7bf218ffea9865ed19524d0da

Observation bc928338-d764-4f8e-b201-698f2db36b69 · outbound

This paper cites Nationality bias in text generation.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Nationality bias in text generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.369665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.917276Z digest=sha256:64e48564716c043e639cde18a77472ea5d5b84056c98524c0c14336e35c00408

Observation 5ff7eb45-af62-47df-b791-e1d12b8744d1 · outbound

This paper cites Biases in large language models: Origins, inventory, and discussion.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Biases in large language models: Origins, inventory, and discussion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.354171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.921954Z digest=sha256:7c48ac37fa4029574ffd60bc960858d424bf32b7ba4db42f9bcdc82b80fbd41d

Observation 847c158c-e188-4133-9580-7cf2bc72b50b · outbound

This paper cites Mechanistic interpretability, variables, and the importance of interpretable bases.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Mechanistic interpretability, variables, and the importance of interpretable bases

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.337615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.928556Z digest=sha256:596ae865717e567e9bd6de82c29e76778107a1949b1196879466b98155f6702c

Observation 9d893a85-de54-4037-a156-6432b71273b4 · outbound

This paper cites Zoom in: An introduction to circuits.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Zoom in: An introduction to circuits

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.934254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.934254Z digest=sha256:c9da87b78c7973f1ef45bc6d1f994be48040802f891576406656304b68b10b98

Observation 135a2cb1-9ab0-45a4-b8d0-a2cede383089 · outbound

This paper cites Gender biases in automatic evaluation metrics for image captioning.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Gender biases in automatic evaluation metrics for image captioning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.310828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.942668Z digest=sha256:71313dd13752711236a33fb52a862697af6ccbf450f5e262264883c93f70d702

Observation 0cf90f9a-f802-4b7f-a6c8-98685c9a04c0 · outbound

This paper cites Language models are unsupervised multitask learners.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Language models are unsupervised multitask learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.948884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.948884Z digest=sha256:cda11b0a4c5f6a95ff2c71184b4e4c4139b4401883e2c4f34c9c95b423dd64da

Observation c7fc6bcb-07b2-4be0-9c34-a6165b7b0261 · outbound

This paper cites Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.953942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.953942Z digest=sha256:0417d4a81bf08f0f389120dfc631e5b11a88b7cf63d8884bf5f573c2057b07d4

Observation 8fede18b-8b4e-4c08-b56d-f4889fe45dbd · outbound

This paper cites Investigating gender bias in large language models through text generation.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Investigating gender bias in large language models through text generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.285328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.958941Z digest=sha256:d4344676d8c555e6dc897e7a0dcba32f4ae5e9c1396acfb27547448af228d527

Observation ef6e9eec-cce5-4f92-9452-88191b5a20b1 · outbound

This paper cites Attribution Patching Outperforms Automated Circuit Discovery.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Attribution Patching Outperforms Automated Circuit Discovery

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.964417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.964417Z digest=sha256:b2fcd4534b87768b2e7a3695b52616dc57f7100cf237b76f1f86229f8c955bbd

Observation 7bb28f1c-b4cf-4562-bbc5-b74d822a9173 · outbound

This paper cites Attribution patching outperforms automated circuit discovery.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Attribution patching outperforms automated circuit discovery

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:29:45.268077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:29:44.969616Z digest=sha256:db471ab86922eb3c5546745ab6f5ca659caaf544f131ff747d76b1d60e259413

Observation 6d1dcac1-8530-40f4-81a1-002bab812bb4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.976406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.976406Z digest=sha256:45d41adfdff18cab72fcc36f16f930f8a8805ff8bc0c80bd401b0e8bed0eaa27

Observation 6e513610-6d58-4c62-a29b-0ed757fc2a36 · outbound

This paper cites Investigating gender bias in language models using causal mediation analysis.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Investigating gender bias in language models using causal mediation analysis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.981245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.981245Z digest=sha256:2c210adb77ccb484874a24a8fdc25bdb11df2b54b966c0041adeaa5ba5832abf

Observation 795ad8e5-0f9c-470e-b00f-ec2eff7fd9fe · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.987123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.987123Z digest=sha256:8d033119573401da4995fd7bc4d7995e091110cf1a7a4ad5e1c69443c36e2cd6

Observation 26a9f4e5-0caa-4643-80c8-0ee51cb33f37 · outbound

This paper cites Neural Network Acceptability Judgments.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Neural Network Acceptability Judgments

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.993184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.993184Z digest=sha256:fa88bcb88e7fca4f834675cfdff0203a8c02e62bcd4e5946443f887bab52cc5e

Observation 79525a86-b6b0-4a39-b3e9-d08605577f50 · outbound

This paper cites Interpretability at scale: Identifying causal mechanisms in alpaca.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Interpretability at scale: Identifying causal mechanisms in alpaca

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.998131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:44.998131Z digest=sha256:ce74ff6caac66292b5a30f9ea9a2752271a9db4fd8efecbe7a15be6e768493ab

Observation 45d1bd76-2e0c-4611-b48c-608a999fa62a · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:45.002961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:29:45.002961Z digest=sha256:9bf0214669febc3e7bcc0e002675d93b0d7b582973419612ec7c136cf57c50db

Pith citing papers

Observation 6a2ef335-0406-48f7-87ce-ccb2a49563dd · inbound

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability cites this paper.

Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:04:06.753097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:04:06.753097Z digest=sha256:5bce0c36985e4de1ff4c22096c915293982f8da9b4046da4fb45c042cc6c5b35

Observation aa71bf9e-5315-4636-9005-727082cadf69 · inbound

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations cites this paper.

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:58.465833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:58.465833Z digest=sha256:323b4b0e33a729d4ada29d08e4783d9b619e36d167134f0609c4ff29ead2c72d

Observation 75533c7d-b5d6-4745-b05d-4ef1be771b61 · inbound

Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models cites this paper.

Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:18:00.320812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:09:28.032659Z digest=sha256:2e83fed242a34dde1d26a20c29bbd51d0fe5a219119f9dfc288d4588e6d60507

Observation d9d1ab2f-48c3-4301-a161-5a866954f5e9 · inbound

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender cites this paper.

GKnow: Measuring the Entanglement of Gender Bias and Factual Gender Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:52:16.728754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T04:50:43.547037Z digest=sha256:d02dce978a1b6b3e698716ede7b391c695b6a2678857bc3321e5f28b54307f3a