Pith. sign in

Paper Citation Record · LEDGER

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

As of 10 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2506.00253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00253 v3

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:20.054591Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact14
  • verified fuzzy2
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 584af120-bfe6-40c1-b645-8947fdc8da08 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.741659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.741659Z digest=sha256:09ebaecffa3489a66441ba7b44c802d60c413785f97bc06aec4a248cef6da16e

Observation 6693e333-a4bc-4cd5-aab8-f4c3287e2506 · outbound

This paper cites Apfelbaum, Michael I.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Apfelbaum, Michael I

Reference 2

Resolution
verified exact
doi, observed 2026-08-07T12:13:21.009857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:13.794504Z digest=sha256:b397b9d732385050ae8523b279222117374333de8588a5573c781f1a14e739ba

Observation 6959ccea-181f-4160-a58e-67650f41031f · outbound

This paper cites Griffiths.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Griffiths

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.852499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.852499Z digest=sha256:98ac402c985e5d3c38e2f3db5212d277c0409bb8dfe2cbc2e78045cca3c58ff9

Observation 79169ae5-5c3e-4610-9b62-d5849c130c93 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:13.935489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:13.935489Z digest=sha256:3ad21d46b7869b6543c1d5e00e9c604c7205be3a6bfe4197c9c05eb8837be181

Observation dee7fa67-5010-44a5-a867-7fc4e4cf00ed · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.003675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.003675Z digest=sha256:7113511bf996d92aa01c69f3a77068dcafa703ab8a68a57bdf8878188f24b4a8

Observation 25500e5a-d219-4d35-a161-615a1926a79d · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mechanistic Interpretability for AI Safety -- A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.094600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.094600Z digest=sha256:379fbcbed5677da6bf54481021da93da28ef2e5aa165f830c8b3eb3f70436903

Observation b9021306-b3f1-4f09-a292-b7889ada3133 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:24.122596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.161233Z digest=sha256:4c726367224a25f3c5170a25d12eca61c1fb02a3a6236f26dd4867bb54ac4b95

Observation ec3a1bb6-f211-4382-a1a5-f523cad9fb58 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.955589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.260481Z digest=sha256:6c5ed695193c8807f3e8ed8de254afd562be348a6a1dde3c5f9d51390401a362

Observation b587f0b4-5934-44b7-83d5-8e94695885a6 · outbound

This paper cites Bryson, and Arvind Narayanan.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Bryson, and Arvind Narayanan

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.350152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.350152Z digest=sha256:b4546bae439b3cdf0cef781d9da7d692851b453d73a2deeb1b2783acb3ba0056

Observation 00cfff6c-03a5-4672-a484-d602f0ba15bd · outbound

This paper cites SelfIE: Self-Interpretation of Large Language Model Embeddings.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SelfIE: Self-Interpretation of Large Language Model Embeddings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.424151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.424151Z digest=sha256:23bcb6dcbef9f1dd19d0da32448272370bf86bd9ca28f4d39f842c128c1d99c9

Observation 38c54000-6257-4f66-ab11-430cc0e4cb18 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-07T12:13:22.867513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.501115Z digest=sha256:68c721ea844a05833f41020df69638b14ed7df15a8ac0a9817a7d5f936360e74

Observation 08e949fc-017d-46d7-a534-d69180fd89ee · outbound

This paper cites Mitigating Social Biases in Language Models through Unlearning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Mitigating Social Biases in Language Models through Unlearning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.600807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.600807Z digest=sha256:8982edd48b13275e2c1018c3e75cd7538c23514d7f73b01312ccc8f2fad6191e

Observation b2146111-85f9-44e2-b22e-f2e110c592eb · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards A Rigorous Science of Interpretable Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.667487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.667487Z digest=sha256:9827482b0865b039c6c9fe6c3c49d893c8544304f98b9e01e6769ee8c73050b3

Observation afe9ad33-06bc-464b-97c0-6a8a5d113ab8 · outbound

This paper cites Eberhardt, Phillip Atiba Goff, Valerie J.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Eberhardt, Phillip Atiba Goff, Valerie J

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.812539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:14.750088Z digest=sha256:4781592879158f89fffde3eace951a054a85c09d3c6771605cbc1239df13f2b3

Observation 58b48d5e-0cb3-401c-a28c-3804aefd0595 · outbound

This paper cites Causal Abstractions of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Causal Abstractions of Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:14.859364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:14.859364Z digest=sha256:a5b0b24a6abebf90c1459f2345b5545f4385b8c81e6dd4600b1dc1b6d6193183

Observation 78865775-a7b6-4d8b-a38d-51be5ae4aba3 · outbound

This paper cites Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:15.219580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:15.219580Z digest=sha256:b96ec182a0ca424deaa1af9f2dd5fb032fbcbccbe072219422a3e44848d3ea91

Observation 4586dff6-61a8-4446-bff3-2bb97f195644 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.064697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.064697Z digest=sha256:484c33b556a0465670b9c35396ad6f535ff49ea5259e47b88814d598e9b453b0

Observation df06e2fa-0b1f-4ae8-969c-f9f447510e4a · outbound

This paper cites Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.597821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.597821Z digest=sha256:85066b003e2fb19f988cd69f5aa0788c7c426a8a9bc7b195f7f835e63d07bf0f

Observation ff1a6bf1-8d9a-4d56-a162-33a4b15fe7ab · outbound

This paper cites Greenwald and Mahzarin R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald and Mahzarin R

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.659024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:16.686678Z digest=sha256:b1a22ca1ce8659c8f7093371cfb0560161ade0457e0089d3d106a9d11a6f33af

Observation 35e296b4-9900-4eca-ad87-17e749474d98 · outbound

This paper cites Greenwald, Debbie E.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Greenwald, Debbie E

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.789954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.789954Z digest=sha256:4c8c28b846bbed7381db059ba6d95e83347988f75ebd46b51fc8fbe5135ab68b

Observation 224ed99f-e562-4cb1-b849-0c55c27b0dff · outbound

This paper cites Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:16.945425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:16.945425Z digest=sha256:2f09816a36027038b26463dbf15f155207b4fdb3ab0ef3f2cf0f50fc8c2ee809

Observation d49ceaf5-1daa-4fd0-a90d-dba6ef10d58f · outbound

This paper cites Language Models Represent Space and Time.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Represent Space and Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.067308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.067308Z digest=sha256:e30d5655fdb950966a9c7e9a52e3b040765d9d73e0ce743a62ac623ca5a13313

Observation b85fd806-9985-45e6-81e9-159e2ab2f9e0 · outbound

This paper cites How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.160012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.160012Z digest=sha256:6703c4c4eaaba078f914c361e39df65a15341f017f936f23e470962f76614e29

Observation d2637206-1289-405a-96cc-507c0d022e51 · outbound

This paper cites How to use and interpret activation patching.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race How to use and interpret activation patching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.249944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.249944Z digest=sha256:87bf7a32d66a1bf0e8bf36c8534669206bacc8aeeb6c2fac07e72ba4a59ac0b5

Observation b4570df3-89e4-456c-a36b-1130f1215455 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.549384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.354469Z digest=sha256:ff1aad6a8f1b1370864ba036a589b1fc731da47550fa59984e5a7d08f71ca808

Observation 3acbc310-04b9-4d3a-8ac8-f5ef511f3555 · outbound

This paper cites Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.479369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.479369Z digest=sha256:4c0d2027eb7acc5e756d34921dcfc107508fff7306a9219f4cff4ecd7284bbe4

Observation 480c2b11-832e-45c7-b6fd-5f6b41f0b59d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race LoRA: Low-Rank Adaptation of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.522209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.522209Z digest=sha256:341055b1214de24d09d007cf914cee07101aa81459793dcc545742c75048816c

Observation 243aac18-e9b7-4429-92a4-baab3b8899e9 · outbound

This paper cites Auxiliary task demands mask the capabilities of smaller language models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Auxiliary task demands mask the capabilities of smaller language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.581753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.581753Z digest=sha256:e501964214a15c17f41b68dd6824cb5a4c3413116efb8512f8bd93e780f1ab71

Observation 37a768e2-1405-43ac-9b6f-612938c3823b · outbound

This paper cites Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV).

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.616483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.616483Z digest=sha256:bae1a4f6e4184d05edafdb39430545476f928753fdc0ed43987c4ce01a4bbdd9

Observation 334c5b38-8c20-4ff2-9902-768826231db7 · outbound

This paper cites Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:22.496299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.673015Z digest=sha256:a042be83fa55ddae54d5755db26874707a6dc3270e4f9950246cca99c99e5d0a

Observation e8c1a24c-31d4-49f9-8c43-df6bbe522c10 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.735447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.735447Z digest=sha256:95762c08b86355b7ec8dde119bd269403060974218fef9d0b5db56c025d68f8b

Observation c8fc1369-dbc1-4c58-9090-afafa544b429 · outbound

This paper cites Levinson, Huajian Cai, and Danielle Young.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Levinson, Huajian Cai, and Danielle Young

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.767619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:17.789820Z digest=sha256:3e1d5cb1a5265eef6ad2a9658afa405a6301ed0c19165b8d31e23d0463705842

Observation 66717e8c-5316-418f-89fe-dd5ecaf976b1 · outbound

This paper cites Inference-Time Intervention: Eliciting Truthful Answers from a Language Model.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.840290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.840290Z digest=sha256:cfae886a56ff0a0710eaa76ade22f3c1f539e06f6bcbfc8c2aa37d36fca6a346

Observation a41f2c33-8fd2-409f-bcd0-ee0654ea817c · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.876984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.876984Z digest=sha256:0214c98f17c37982be106eec9c50dd4a92bcb2a399e2bdd0664f988337cc8b2d

Observation 2a9d59b5-2f6e-49be-a7f3-b9a841bab127 · outbound

This paper cites Label Supervised LLaMA Finetuning.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Label Supervised LLaMA Finetuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.913243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.913243Z digest=sha256:7a47dfed65eda9bda6b9cecdc307bd4f9a515c04e429eb60167a9c703ec4bae1

Observation 862d51d8-9230-4423-a187-bc9d42e4369f · outbound

This paper cites Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.953463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.953463Z digest=sha256:98d74926772ef4024f6c1992d83376bc67e95cea5c9cd7eab878ffda165cd14f

Observation 2a1acba3-eaf7-4cb5-be78-f6d35f2f4183 · outbound

This paper cites Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Assessing Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:17.999979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:17.999979Z digest=sha256:2ce9f22cca432581403cbe28c7c074033292d938a88bcb353f1ba285556e42fb

Observation 1e5db672-887f-4e75-befc-cd4d69882a92 · outbound

This paper cites Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.050267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.050267Z digest=sha256:fe3a80afd17ccf84bd219045a1d3d3aa7acef99f53f0b28a9765dd1ad851fd54

Observation e298aa2c-a53b-4891-9675-da7ad6e89c15 · outbound

This paper cites The Llama 3 Herd of Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.104681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.104681Z digest=sha256:4d16138030c06585e1a1be57333f764111fbe999807234603c08d80a5350df99

Observation 4a7b237b-bf63-47e0-ad39-b82784fff478 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.154269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.154269Z digest=sha256:2e383108be62bfdc92bd1bd8dc0c931786fde8f316a4b61c90e0f2b1338ad9c1

Observation 147bb845-edff-422e-92ce-19d5829b3ec8 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Locating and Editing Factual Associations in GPT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.187628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.187628Z digest=sha256:da0aa2da9e5818cdd74afbcb44e6de61dc85215a041a7c97424ac85654a22d7c

Observation aebf94c6-0dcc-4975-9a8d-b3ee07eb6fc7 · outbound

This paper cites Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.204599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.204599Z digest=sha256:1a8fed158874f397c69f61cf9262b5c863f1dafba71edda8326e1770c268bb2e

Observation cc38fca4-fe45-431c-9c57-0e03e2cae629 · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Progress measures for grokking via mechanistic interpretability

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.242414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.242414Z digest=sha256:18a06874867b9c81b491dbc7f481400b92da3fa2cdd53d9213f323847213e016

Observation c2dfeb36-b961-4fc3-8601-ecc849ada23a · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.396510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.308267Z digest=sha256:76a1c03f85a0d09beb257cc1f8fcc50e9eb999e6b6ea5756d5ecb399d0cf97dd

Observation a4e02824-91e2-44eb-8315-2d05e6bbe815 · outbound

This paper cites Norton, Samuel R.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Norton, Samuel R

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-07T12:13:22.193037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.353714Z digest=sha256:d73274cda0001c226677142b34c475d0a913c07ee5dc2fcb0b2ff6c737cda8b0

Observation d5d56511-f8a6-4701-b478-93e5dfc9d302 · outbound

This paper cites Nosek, Anthony G.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Nosek, Anthony G

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:23.606008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.428035Z digest=sha256:f8dc1936d937ea2971d28928365c1fe8caac2ac2350483ce56bdb556f57102d9

Observation 663161dd-2dd9-4ce3-b2c0-a9d32f946e9e · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.461680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.453442Z digest=sha256:71ce31462a1d523eed1880374c610760e93b7996f717be0db7b40bd69a16b456

Observation d7e20d09-c0fc-438d-8b05-7fd6ebaa2492 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.346962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.494148Z digest=sha256:d439ed284650c1f31bbae073fb83d836f6ec44f6b702b5a9ce016d8268edda9e

Observation f53dca34-7fcd-425f-add1-d764bb3a832f · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Llama 2 via Contrastive Activation Addition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.522808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.522808Z digest=sha256:1532da3dcf4e3e20f20b730e4e07d61db467da96d3b2e9bc58a555bf0f3803be

Observation fa03f3aa-6535-4e1f-814a-d6d57c3fa12a · outbound

This paper cites BBQ: A Hand-Built Bias Benchmark for Question Answering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race BBQ: A Hand-Built Bias Benchmark for Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.554467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.554467Z digest=sha256:7de15fbdde27b476c4b9c077f3e61915f90c2d7e201c4a0fe526a9f068809b36

Observation 7f9206dc-82bb-4d97-bbb5-fddde5ac9efc · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 51

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.248908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.586340Z digest=sha256:a3ed3c32a86e40351799f0adaf90db28c72ce8e3b698235bfbc74bcb32d4a8c3

Observation d18a9a1e-ae54-49ad-95ad-42e3b8f0955e · outbound

This paper cites Interpreting Bias in Large Language Models: A Feature-Based Approach.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpreting Bias in Large Language Models: A Feature-Based Approach

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.933327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.626848Z digest=sha256:d25bd1a8bc9091c3982082398bfb3dc6e5e3da34acb0c92fd0154f2502f009c6

Observation bd3f5835-517f-4287-8a21-971b9f3e7c35 · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.195270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.673799Z digest=sha256:f2f074d122b8db04cf7812165c24eb1712b7cf6e7444652bc87a41c0784c54f8

Observation 2d48be1c-cd16-4ea7-a872-ed5fd0e79589 · outbound

This paper cites IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.714338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.714338Z digest=sha256:952d3a3491a58b47aab746d6449ce073084c9b3d215077f1026c791117ebe7bf

Observation 9d141c4d-440b-4d20-8b30-a8982e54cb9d · outbound

This paper cites Efficient RLHF: Reducing the Memory Usage of PPO.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Efficient RLHF: Reducing the Memory Usage of PPO

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.743392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.743392Z digest=sha256:a1096b3433cbb50577cb2526d9c9de05c599ca6354907828d35d40812d492bb9

Observation a58d3885-7902-467a-b295-3112b6003fad · outbound

This paper cites Parameter Efficient Reinforcement Learning from Human Feedback.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Parameter Efficient Reinforcement Learning from Human Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.772325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.772325Z digest=sha256:567accfb3cf35c2416a8a5cad4cf1469e76fa6c00fa23fe2c2bfca3653245c2d

Observation e5212170-b945-4119-b102-0a9a94a2f5bc · outbound

This paper cites Stevens, Victoria C.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Stevens, Victoria C

Reference 57

Resolution
verified exact
doi, observed 2026-08-07T12:13:20.147034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.804298Z digest=sha256:91aada4c446a28a2a1942c8bc4b41b79b6b34c6ff76ca55c44d9ab67cc94f93f

Observation 9657af22-7b98-451f-a8a8-63c6060afb66 · outbound

This paper cites Improving Instruction-Following in Language Models through Activation Steering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Improving Instruction-Following in Language Models through Activation Steering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.854878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.854878Z digest=sha256:c6c854bce1903f8877cbb16f24e8c0b098e914d33aa8f41a1ef2b8a216fbc243

Observation d67a5ee7-ab89-4cbd-8a08-ebfbdb2eecfb · outbound

This paper cites Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.882947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.882947Z digest=sha256:26d7314a92fe0aed68d80a2e757045de39ddeeaf54a454141ada1c53e2fb64c6

Observation 38452e32-1802-4a5e-b743-99108b95aee0 · outbound

This paper cites SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race SuryaKiran at MEDIQA-Sum 2023: Leveraging LoRA for Clinical Dialogue Summarization

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.647590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:18.924741Z digest=sha256:3a928ea78830ea1a5126d84fb47c9b70e0cd3036ccaec5efd028e81d80f626f2

Observation 68ee41b9-5b47-418d-b855-3d6afbe5dc61 · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:18.979922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:18.979922Z digest=sha256:97fd1e3f06931c238b0a37d1e2e942448a54da094b29e49ce7f2e1107857d753

Observation 2e7df569-9d9b-41c6-9fa8-1a2fe05f9e2e · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.036340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.036340Z digest=sha256:803d2487a5808277161112e8b2a8f8045cfe60b6072d9316e02c3490316e023a

Observation 33d965b8-0bd6-42eb-8d4c-014e25b9a186 · outbound

This paper cites Steering Language Models With Activation Engineering.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Steering Language Models With Activation Engineering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.106538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.106538Z digest=sha256:8b7d35f43b9df7b6ddb71cb5bfb976036a8c15d4433a2213ae2c2e81b5291b81

Observation 2274232a-1bde-4b51-a43f-f0d4ec7d6e0f · outbound

This paper cites an unresolved cited work.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:23.082467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.178953Z digest=sha256:bc9e6820fddb963811555611fd91bd8e8d419c0eaed63889a70bc7b8fa80cae7

Observation 5724ed5a-dc59-49c0-9876-5279146e77aa · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.236371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.236371Z digest=sha256:3b7871b34b7fe79883742ad40eea6545f9bbc23a6037682fb0c57143eeb8d2f4

Observation 25085deb-1076-4066-a0b5-1f3681d0afe1 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.352932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.352932Z digest=sha256:4a5782ececfd2ae0344df5e1471abdc51543282e516bbf15ee1e6eabc62b2188

Observation a9b299b5-3400-43f7-846d-22f590859600 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Jailbroken: How Does LLM Safety Training Fail?

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.432193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.432193Z digest=sha256:e68141fb0e0d702f7645009a05385c6b0b7397df37d9d3614afc5cc6a11d2684

Observation 5e43b040-e749-4864-b1e7-84ea0d259d51 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Fundamental Limitations of Alignment in Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.496083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.496083Z digest=sha256:9a57b39e1d5d2f6d8a667f86bcde274048793aeb1eacef4d8dac00ec13024f54

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · outbound

This paper cites AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:bcacf8df79f32d20ee213395c7f0fd16a5ec96c2e4a7553e6747c639c607746e

Observation 5864e4db-8df3-4dce-830e-4957a71edba1 · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.638500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.638500Z digest=sha256:03a10c5799fd4552594cf27715110e04e8b251106860a37c7166fd38969ec31c

Observation 2a13aa06-7e8b-4911-9eec-797247b2bfa4 · outbound

This paper cites AutoRE: Document-Level Relation Extraction with Large Language Models.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AutoRE: Document-Level Relation Extraction with Large Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.360912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.714890Z digest=sha256:353153ee1c08321d041f0cda86ca2dab7b616fc4f78033c1b84a4d3224760644

Observation 94cf689b-5052-4457-8597-023c50383804 · outbound

This paper cites Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.780988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.780988Z digest=sha256:930501681af1183b19360949e8c09e6d72d395935389967a82fc756d201a9f07

Observation 60034b32-2269-4448-a321-ae31c92f7089 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.843762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.843762Z digest=sha256:c6fc76d3586d7bc0dce05bd66767135031e5c4bf2de1e05cc6a64de99205ad2c

Observation 6a303adf-c8b0-47a8-90e1-7d81d48f201d · outbound

This paper cites DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race DialogueLLM: Context and Emotion Knowledge-Tuned Large Language Models for Emotion Recognition in Conversations

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:21.217774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T12:13:19.901587Z digest=sha256:1530765f309a4942824f4eaf85c1bccc3380a510de3572040ee4bf693b0ab890

Observation 24a1d6c1-3256-4b83-bf4b-bf345a072fad · outbound

This paper cites The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural Networks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.962960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.962960Z digest=sha256:3a11e2adc0a72ed4ae83148127828004d7da719ccab7121e069b2ec269c354af

Observation c4d24a5c-c423-4b35-bcbe-afa331c8565e · outbound

This paper cites online" 'onlinestring :=.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race online" 'onlinestring :=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.045313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.045313Z digest=sha256:8fce3a85c5728868ecbc2d3c4799309fd93fb13d43ca9227e21c647fa1d2351a

Observation 76c487d5-46d7-409a-9876-765a1bab2955 · outbound

This paper cites write newline.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race write newline

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:20.054591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:20.054591Z digest=sha256:0f7b44033be9a3930ce6953953c4e97cb699291abb065cb4e680c447f35096c3

Pith citing papers

No inbound Pith citation observations are available.