Pith. sign in

Paper Citation Record · LEDGER

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2506.00253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00253 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:20.054591Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact14
  • verified fuzzy2
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 584af120-bfe6-40c1-b645-8947fdc8da08 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.741659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.741659Z digest=sha256:6ffde77bdbebc054531a0ee8eaf5a2476739eff58d3dc63fba65fa7b52a23471

Observation 6693e333-a4bc-4cd5-aab8-f4c3287e2506 · outbound

This paper cites Apfelbaum, Michael I.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Apfelbaum, Michael I

Reference 2

Resolution
verified exact
doi, observed 2026-08-07T12:13:21.009857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:13.794504Z digest=sha256:cfd7db1264d3388ed7cd618d0daa8189962187471e89e0cda6dc8680573c4941

Observation 6959ccea-181f-4160-a58e-67650f41031f · outbound

This paper cites Griffiths.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Griffiths

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.852499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.852499Z digest=sha256:98ac402c985e5d3c38e2f3db5212d277c0409bb8dfe2cbc2e78045cca3c58ff9

Observation 79169ae5-5c3e-4610-9b62-d5849c130c93 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.935489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.935489Z digest=sha256:3ad21d46b7869b6543c1d5e00e9c604c7205be3a6bfe4197c9c05eb8837be181

Observation dee7fa67-5010-44a5-a867-7fc4e4cf00ed · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.003675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.003675Z digest=sha256:3fedb1f9b258806d2acbc1c32e73888a71f881d9f49664c12ab51056b6f262e6

Observation 25500e5a-d219-4d35-a161-615a1926a79d · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mechanistic Interpretability for AI Safety -- A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.094600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.094600Z digest=sha256:379fbcbed5677da6bf54481021da93da28ef2e5aa165f830c8b3eb3f70436903

Observation b9021306-b3f1-4f09-a292-b7889ada3133 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:24.122596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.161233Z digest=sha256:06daf7e7af3a448a661caec03c14bf202c177512bf65b3953e88e083b0f57154

Observation ec3a1bb6-f211-4382-a1a5-f523cad9fb58 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.955589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.260481Z digest=sha256:76d8524b5a60bb30510a7e8b06caf61fd897e3536bd2ce04c0a7f45251ada680

Observation b587f0b4-5934-44b7-83d5-8e94695885a6 · outbound

This paper cites Bryson, and Arvind Narayanan.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Bryson, and Arvind Narayanan

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.350152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.350152Z digest=sha256:b4546bae439b3cdf0cef781d9da7d692851b453d73a2deeb1b2783acb3ba0056

Observation 00cfff6c-03a5-4672-a484-d602f0ba15bd · outbound

This paper cites SelfIE: Self-Interpretation of Large Language Model Embeddings.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SelfIE: Self-Interpretation of Large Language Model Embeddings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.424151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.424151Z digest=sha256:48618d75085780d1f4ce1907f0fd0677a57befdf0540336c0bc9bd555656aecd

Observation 38c54000-6257-4f66-ab11-430cc0e4cb18 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-07T12:13:22.867513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.501115Z digest=sha256:b3cf6cf368a2f5411f81b269664a8d90d6dcf1ee16efee85d67f14e0be25a89e

Observation 08e949fc-017d-46d7-a534-d69180fd89ee · outbound

This paper cites Mitigating Social Biases in Language Models through Unlearning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mitigating Social Biases in Language Models through Unlearning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.600807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.600807Z digest=sha256:363381d9f1be329f2bb65a3286f66c14e8fd64a339e2e821de80d422075cf144

Observation b2146111-85f9-44e2-b22e-f2e110c592eb · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards A Rigorous Science of Interpretable Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.667487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.667487Z digest=sha256:97e15301e4997cee23f0ce13901a309b65848fbada3cf1a0c1639f1998f3ee7d

Observation afe9ad33-06bc-464b-97c0-6a8a5d113ab8 · outbound

This paper cites Eberhardt, Phillip Atiba Goff, Valerie J.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eberhardt, Phillip Atiba Goff, Valerie J

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.812539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.750088Z digest=sha256:9c819322144304758b043f81f6e4c8fd0f21bff45510c1a69336ee06dac78c9d

Observation 58b48d5e-0cb3-401c-a28c-3804aefd0595 · outbound

This paper cites Causal Abstractions of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Causal Abstractions of Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.859364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.859364Z digest=sha256:aab4e9a14170fabe1f0cd736579540739ea0c289cad7f79843036489e8f4240e

Observation 78865775-a7b6-4d8b-a38d-51be5ae4aba3 · outbound

This paper cites Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:15.219580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:15.219580Z digest=sha256:dc68e3e79ae0daf4e7f1c6f50e9a57e8319bc7d2dd0159009006f17ae57eef02

Observation 4586dff6-61a8-4446-bff3-2bb97f195644 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.064697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.064697Z digest=sha256:03acfd37aa507de71a061d81a3ce42a2b255500e05a5778c2e33a4b749e59761

Observation df06e2fa-0b1f-4ae8-969c-f9f447510e4a · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.597821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.597821Z digest=sha256:483ea7c53f0cfec0101ead652c75d5c01fdfc6034e6baf4235095399ae4bc7de

Observation ff1a6bf1-8d9a-4d56-a162-33a4b15fe7ab · outbound

This paper cites Greenwald and Mahzarin R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald and Mahzarin R

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.659024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:16.686678Z digest=sha256:538b0ef1ede52d07d810cdda9f5685a46beb54bdcb707bf82324ad6f9ea3b008

Observation 35e296b4-9900-4eca-ad87-17e749474d98 · outbound

This paper cites Greenwald, Debbie E.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald, Debbie E

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.789954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.789954Z digest=sha256:4c8c28b846bbed7381db059ba6d95e83347988f75ebd46b51fc8fbe5135ab68b

Observation 224ed99f-e562-4cb1-b849-0c55c27b0dff · outbound

This paper cites Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.945425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.945425Z digest=sha256:85efa76e2250cbe88c91e969dbf2d44efa8b15b1f40cfe45245ae6637cfd9445

Observation d49ceaf5-1daa-4fd0-a90d-dba6ef10d58f · outbound

This paper cites Language Models Represent Space and Time.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Represent Space and Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.067308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.067308Z digest=sha256:9f2e2aa6d1156dc4d1c5d1681a98422344e5be8ad3d8c404aff7a3c9aabc4dff

Observation b85fd806-9985-45e6-81e9-159e2ab2f9e0 · outbound

This paper cites How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.160012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.160012Z digest=sha256:02f54394ed7f517b6c9401d49b5a01f3070f4ff2827c5d9dd63d3e5dc8eb08b8

Observation d2637206-1289-405a-96cc-507c0d022e51 · outbound

This paper cites How to use and interpret activation patching.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How to use and interpret activation patching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.249944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.249944Z digest=sha256:5fc4ef3dbcc47f659205cf9c793837b6ea5c9e7739fb78f495f55ece1eb7d46f

Observation b4570df3-89e4-456c-a36b-1130f1215455 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.549384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.354469Z digest=sha256:a4e28e417e7f5ac117e61466e98d663c3b17d3101867606a4fa18fc3269c9468

Observation 3acbc310-04b9-4d3a-8ac8-f5ef511f3555 · outbound

This paper cites Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.479369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.479369Z digest=sha256:ff1e22178723e5b199ff2d8622b8ff2c91d13070b958c1b07ba32bb1e9f94600

Observation 480c2b11-832e-45c7-b6fd-5f6b41f0b59d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race LoRA: Low-Rank Adaptation of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.522209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.522209Z digest=sha256:fa537d22d36563be7d9fbfe5d3726e4c468eed186aec4519e9d006dd3d7713c2

Observation 243aac18-e9b7-4429-92a4-baab3b8899e9 · outbound

This paper cites Auxiliary task demands mask the capabilities of smaller language models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Auxiliary task demands mask the capabilities of smaller language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.581753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.581753Z digest=sha256:99f673c3d63370b5098284cb93f06b961dcec7630dbbcccc1ee74e588e1aa27b

Observation 37a768e2-1405-43ac-9b6f-612938c3823b · outbound

This paper cites Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV).

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.616483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.616483Z digest=sha256:bae1a4f6e4184d05edafdb39430545476f928753fdc0ed43987c4ce01a4bbdd9

Observation 334c5b38-8c20-4ff2-9902-768826231db7 · outbound

This paper cites Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:22.496299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.673015Z digest=sha256:b468fb56e830dd3cac8f7b8e40a0a2e3d7e94d4c1860d56c64b83b115f8d2cb9

Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.735447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.735447Z digest=sha256:19cdaa2d214590e9e44e4495ce87bfd77abd20ce7a215895c9154bcad7c92da8

Observation c8fc1369-dbc1-4c58-9090-afafa544b429 · outbound

This paper cites Levinson, Huajian Cai, and Danielle Young.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Levinson, Huajian Cai, and Danielle Young

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.767619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.789820Z digest=sha256:04e2a61f64eab0fc639a312074f64c39bd79a1408015f68fcc55fe87f8169588

Observation 66717e8c-5316-418f-89fe-dd5ecaf976b1 · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.840290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.840290Z digest=sha256:5a11f969acf54a4bb277b5039ab2560bb7ec7c7c691a8b6a559e7bf99d08f5dd

Observation a41f2c33-8fd2-409f-bcd0-ee0654ea817c · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.876984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.876984Z digest=sha256:843292dd89030de7cec3d230e76c62b5399cf3d978dd9c1d844576a1f800d542

Observation 2a9d59b5-2f6e-49be-a7f3-b9a841bab127 · outbound

This paper cites Label Supervised LLaMA Finetuning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Label Supervised LLaMA Finetuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.913243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.913243Z digest=sha256:9fbd120e2e6d5f1b25d8910d8cc3227f730bfb62c54243e24c34fb39030da064

Observation 862d51d8-9230-4423-a187-bc9d42e4369f · outbound

This paper cites Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.953463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.953463Z digest=sha256:a1de8091a704e4993512d8051042f515a1ea5e613f073f4f2c2e1cb7d9be0d8a

Observation 2a1acba3-eaf7-4cb5-be78-f6d35f2f4183 · outbound

This paper cites Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.999979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.999979Z digest=sha256:6968d29ea085e6042990b89ec9e2be332dadd3590dbf268edb0a20c7f4a3c117

Observation 1e5db672-887f-4e75-befc-cd4d69882a92 · outbound

This paper cites Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.050267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.050267Z digest=sha256:770ad93587a0713ff5369271c5d5cd4e31b386e0cdd7e90f87f61019f7d2fdad

Observation e298aa2c-a53b-4891-9675-da7ad6e89c15 · outbound

This paper cites The Llama 3 Herd of Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.104681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.104681Z digest=sha256:0baa22f5e03cc413e3574704f5d80ae326433368aea410b8e23d6c8f4504ce8f

Observation 4a7b237b-bf63-47e0-ad39-b82784fff478 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.154269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.154269Z digest=sha256:8219ebdeee3cc1d493603365a88dfaaff18d13864b678068bf8f5b601500b84d

Observation 147bb845-edff-422e-92ce-19d5829b3ec8 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Locating and Editing Factual Associations in GPT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.187628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.187628Z digest=sha256:ff5c9c5a651d9f7421ce4199e22e7acf718d9bee9093714e324aee098ad32c19

Observation aebf94c6-0dcc-4975-9a8d-b3ee07eb6fc7 · outbound

This paper cites Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.204599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.204599Z digest=sha256:ec46ed9719f440cbaf5538633a48d8a646d684a44eac54d01a78230f328a6825

Observation cc38fca4-fe45-431c-9c57-0e03e2cae629 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Progress measures for grokking via mechanistic interpretability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.242414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.242414Z digest=sha256:18a06874867b9c81b491dbc7f481400b92da3fa2cdd53d9213f323847213e016

Observation c2dfeb36-b961-4fc3-8601-ecc849ada23a · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.396510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.308267Z digest=sha256:fef80bedf352530fd5cd55bc84b47f0c301f5469c41871c4304cab9596303168

Observation a4e02824-91e2-44eb-8315-2d05e6bbe815 · outbound

This paper cites Norton, Samuel R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Norton, Samuel R

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-07T12:13:22.193037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.353714Z digest=sha256:1497d68f8eaa6a4ce09d22411072f14bf345d4d92f0d8c801619cfbbe580da4c

Observation d5d56511-f8a6-4701-b478-93e5dfc9d302 · outbound

This paper cites Nosek, Anthony G.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Nosek, Anthony G

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.606008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.428035Z digest=sha256:9f6036fa69aeb3a4eaeeda1d050440f8c080ce0f6bd4c6317e79324b389ad40d

Observation 663161dd-2dd9-4ce3-b2c0-a9d32f946e9e · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.461680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.453442Z digest=sha256:7d99413c8e057299a00117673dce2b1877930bbf77d02e09559612081b6d25bc

Observation d7e20d09-c0fc-438d-8b05-7fd6ebaa2492 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.346962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.494148Z digest=sha256:7f62950b7cb8482b2105e4e6c653f8a03ed94ebeb98e0150a8cdee26e4c155d7

Observation f53dca34-7fcd-425f-add1-d764bb3a832f · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Llama 2 via Contrastive Activation Addition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.522808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.522808Z digest=sha256:1532da3dcf4e3e20f20b730e4e07d61db467da96d3b2e9bc58a555bf0f3803be

Observation fa03f3aa-6535-4e1f-814a-d6d57c3fa12a · outbound

This paper cites BBQ: A Hand-Built Bias Benchmark for Question Answering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race BBQ: A Hand-Built Bias Benchmark for Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.554467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.554467Z digest=sha256:7de15fbdde27b476c4b9c077f3e61915f90c2d7e201c4a0fe526a9f068809b36

Observation 7f9206dc-82bb-4d97-bbb5-fddde5ac9efc · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 51

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.248908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.586340Z digest=sha256:889ff211b7503f2991150ba933446ed334ee12633807b51a4e823886650b8e75

Observation d18a9a1e-ae54-49ad-95ad-42e3b8f0955e · outbound

This paper cites Interpreting Bias in Large Language Models: A Feature-Based Approach.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpreting Bias in Large Language Models: A Feature-Based Approach

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.933327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.626848Z digest=sha256:cc87c086af0d2b6f14952eec9f23a162d43b2f0ebcb7f30abaab9cee708b3e1c

Observation bd3f5835-517f-4287-8a21-971b9f3e7c35 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.195270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.673799Z digest=sha256:0fe6cb497d8481ec6b7e4657cbda5905d052715182f7e141127d13aa56e1cb75

Observation 2d48be1c-cd16-4ea7-a872-ed5fd0e79589 · outbound

This paper cites IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.714338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.714338Z digest=sha256:1e985f4a5bfeb2971e22c9c59560279c2578a1d1df7f94b26c9dec29cd24341e

Observation 9d141c4d-440b-4d20-8b30-a8982e54cb9d · outbound

This paper cites Efficient RLHF: Reducing the Memory Usage of PPO.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Efficient RLHF: Reducing the Memory Usage of PPO

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.743392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.743392Z digest=sha256:621809c3253f7ea78c7e163a8d8386a67caa00b0640135da1b417e9784b9cef8

Observation a58d3885-7902-467a-b295-3112b6003fad · outbound

This paper cites Parameter Efficient Reinforcement Learning from Human Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter Efficient Reinforcement Learning from Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.772325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.772325Z digest=sha256:6c2a9fd3133c0bb7e3c987451bff62ced1dffe895e37f81117e12fe7d2080011

Observation e5212170-b945-4119-b102-0a9a94a2f5bc · outbound

This paper cites Stevens, Victoria C.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Stevens, Victoria C

Reference 57

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.147034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.804298Z digest=sha256:e20ceac91fdd398b6b9c0a248d73bf941aadfb7aaac77f12bdec28a6b31d334d

Observation 9657af22-7b98-451f-a8a8-63c6060afb66 · outbound

This paper cites Improving Instruction-Following in Language Models through Activation Steering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Improving Instruction-Following in Language Models through Activation Steering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.854878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.854878Z digest=sha256:ad1a457114ca908267e336ac98abeeb870fa156586e659ee93935e253420c8d2

Observation d67a5ee7-ab89-4cbd-8a08-ebfbdb2eecfb · outbound

This paper cites Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.882947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.882947Z digest=sha256:cabe846c4c8bc481de7fc4b33023fe3c78dfce3b160dcf2d456285659dd141ff

Observation 38452e32-1802-4a5e-b743-99108b95aee0 · outbound

This paper cites SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.647590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.924741Z digest=sha256:396a1374799329f14ab8144e5d55214de9e436b58ae57f1bec736f6fcbaf9512

Observation 68ee41b9-5b47-418d-b855-3d6afbe5dc61 · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.979922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.979922Z digest=sha256:f423d8cdf1422277d2432a5e3f0e8544c4c21f1fe39e557a818286ad62bf06bf

Observation 2e7df569-9d9b-41c6-9fa8-1a2fe05f9e2e · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.036340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.036340Z digest=sha256:f7a808b9950c8e8c018f64825d94f22d3fc16250e33cebe75f4c8be8314e0ef0

Observation 33d965b8-0bd6-42eb-8d4c-014e25b9a186 · outbound

This paper cites Steering Language Models With Activation Engineering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Language Models With Activation Engineering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.106538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.106538Z digest=sha256:8b7d35f43b9df7b6ddb71cb5bfb976036a8c15d4433a2213ae2c2e81b5291b81

Observation 2274232a-1bde-4b51-a43f-f0d4ec7d6e0f · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.082467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.178953Z digest=sha256:037ccbcef18e499b193bcffa5273d59ee16d98180b1f827c248314fc405beb40

Observation 5724ed5a-dc59-49c0-9876-5279146e77aa · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.236371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.236371Z digest=sha256:a6da662fe5919fc0c21cb12d1d9b56722a22986079f08a4a1de49ea57723620e

Observation 25085deb-1076-4066-a0b5-1f3681d0afe1 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.352932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.352932Z digest=sha256:4a5782ececfd2ae0344df5e1471abdc51543282e516bbf15ee1e6eabc62b2188

Observation a9b299b5-3400-43f7-846d-22f590859600 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Jailbroken: How Does LLM Safety Training Fail?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.432193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.432193Z digest=sha256:e68141fb0e0d702f7645009a05385c6b0b7397df37d9d3614afc5cc6a11d2684

Observation 5e43b040-e749-4864-b1e7-84ea0d259d51 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Fundamental Limitations of Alignment in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.496083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.496083Z digest=sha256:82291b1033bfba4e152fbaae962adf0a1aec4093763099989eb832fff4ee630f

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:5d5b22850876ba07402d86e2ebc45da8a8df966dc04d3800063f227d8a3f78af

Observation 5864e4db-8df3-4dce-830e-4957a71edba1 · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.638500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.638500Z digest=sha256:5aca9c695a5511a9ad8bf6796eefa9bcc9fb4ae9c3da44c717bd5d71ad4edd8f

Observation 2a13aa06-7e8b-4911-9eec-797247b2bfa4 · outbound

This paper cites AutoRE: Document-Level Relation Extraction with Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AutoRE: Document-Level Relation Extraction with Large Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.360912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.714890Z digest=sha256:ca575868bb3d0d5179a6ad68cfbd7e6eb0e3a007d4b50404dca93f1eb6c9bb65

Observation 94cf689b-5052-4457-8597-023c50383804 · outbound

This paper cites Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.780988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.780988Z digest=sha256:a7892c7c54023b3fbe957575daa75be138fd396ea73bd77d236dab846628be9b

Observation 60034b32-2269-4448-a321-ae31c92f7089 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.843762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.843762Z digest=sha256:c6fc76d3586d7bc0dce05bd66767135031e5c4bf2de1e05cc6a64de99205ad2c

Observation 6a303adf-c8b0-47a8-90e1-7d81d48f201d · outbound

This paper cites DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.217774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.901587Z digest=sha256:307651c971ec7df5f828ed906b721a33363ac5f665efd605162e777103a895e5

Observation 24a1d6c1-3256-4b83-bf4b-bf345a072fad · outbound

This paper cites The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.962960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.962960Z digest=sha256:ba62213318b70178315ab1bf9cd1e29282d7bc8222ce401803d9a8ef15621f4b

Observation c4d24a5c-c423-4b35-bcbe-afa331c8565e · outbound

This paper cites online" 'onlinestring :=.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race online" 'onlinestring :=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.045313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.045313Z digest=sha256:8fce3a85c5728868ecbc2d3c4799309fd93fb13d43ca9227e21c647fa1d2351a

Observation 76c487d5-46d7-409a-9876-765a1bab2955 · outbound

This paper cites write newline.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race write newline

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.054591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.054591Z digest=sha256:0f7b44033be9a3930ce6953953c4e97cb699291abb065cb4e680c447f35096c3

Pith citing papers

No inbound Pith citation observations are available.