Pith. sign in

Paper Citation Record · LEDGER

When Bad Data Leads to Good Models

As of 20 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2505.04741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04741 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:25:00.352149Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:53:35.432279Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:07:08.941549Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2f95a0d-c763-4d66-8f47-141e27d29ba2 · outbound

This paper cites write newline.

When Bad Data Leads to Good Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.094915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.094915Z digest=sha256:2f07e144dd55f346c83a190d7f9d64f25fe083d60cc10cc3a86e5c08e7b86881

Observation 9bdaa413-5218-4268-b5ff-2d0900576d25 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

When Bad Data Leads to Good Models Understanding intermediate layers using linear classifier probes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.101334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.101334Z digest=sha256:3ca403d6c09875464a21f58d7bc046df51dc2c0d53514309f3d37c96b23ae3c9

Observation 2100cedd-d90d-4a1f-b140-cbeec05e5e12 · outbound

This paper cites Toxicity of the Commons: Curating Open-Source Pre-Training Data.

When Bad Data Leads to Good Models Toxicity of the Commons: Curating Open-Source Pre-Training Data

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:25:00.868858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.107667Z digest=sha256:5d916909729cb01b0c4cd96b83377971ed063c43af8f7161db450b18fb988a70

Observation f767017b-facc-44fd-86f7-3a7948583d22 · outbound

This paper cites Linear algebraic structure of word senses, with applications to polysemy.

When Bad Data Leads to Good Models Linear algebraic structure of word senses, with applications to polysemy

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.175216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.112803Z digest=sha256:80ca102a2b5753a68c73c0303a551456dbb6fcd2f302b00bfa441be3365246d4

Observation fb30460a-59c5-40c3-afbe-9aa3508beb6c · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances.

When Bad Data Leads to Good Models Probing classifiers: Promises, shortcomings, and advances

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.159363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.117702Z digest=sha256:e84a5e0685ee2e2575361ee6f2052ae9ff02d6a7d4c2f8133fc1bd631fde1f2b

Observation 9c253f6a-d6c7-409c-b366-ed71761c567f · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

When Bad Data Leads to Good Models Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.122413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.122413Z digest=sha256:e21ea01068c2a04f0b08ecd7312756162e15b5d3d5886e987c187bc590dda894

Observation 1b6646a1-61e8-4c26-881e-672a32d40197 · outbound

This paper cites The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models.

When Bad Data Leads to Good Models The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.127552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.127552Z digest=sha256:190523eff58213bb7deb82e3bce42780c9da24b3ccfe3cab080207fc24e06a7b

Observation 2caa96ed-3b17-4a01-ac18-57aa299bb531 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

When Bad Data Leads to Good Models UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.133124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.133124Z digest=sha256:c788ba71dfd9326a6852f9888515005c676787666e88a030e423b139d29c291d

Observation a2a48bba-7454-4da2-9525-77d95eb0b4ec · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

When Bad Data Leads to Good Models Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.138738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.138738Z digest=sha256:f9b1c59ddb51d73977853b0a7dc12cfd90018c84a9369379ec2f1af00410911b

Observation fa169dbb-ccc3-4820-964d-67a11a660243 · outbound

This paper cites Plug and Play Language Models: A Simple Approach to Controlled Text Generation.

When Bad Data Leads to Good Models Plug and Play Language Models: A Simple Approach to Controlled Text Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.143735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.143735Z digest=sha256:1b66bc742c3bbd67b66879e8ae21850df3d3ecfebb2746e8ab90b28350532273

Observation 4e032bdb-b0ca-41d5-bf95-6951b59003c9 · outbound

This paper cites Toy Models of Superposition.

When Bad Data Leads to Good Models Toy Models of Superposition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.150012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.150012Z digest=sha256:73ea907e45cca6f2d95b4f4acb3aef533d41edb8b791c50c178a4d5669588bfb

Observation 54f300f0-9d0a-4c71-a5cb-a87b170d8f63 · outbound

This paper cites an unresolved cited work.

When Bad Data Leads to Good Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.154504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.154504Z digest=sha256:d9ec55681fb71da6e9fa9dbad91afe5a3f842d3c4e1299c6138bb064e91a6d65

Observation 6961c34a-9882-4e07-b1d2-36dfe273bbc1 · outbound

This paper cites Openwebtext corpus.

When Bad Data Leads to Good Models Openwebtext corpus

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.159644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.159644Z digest=sha256:38b16e3a98048b6f3e2737b0a2796afa53ded61e5612ac5a680eb1d05b3ba351

Observation 9afb00c9-3290-4a60-8934-0424c6d12ea2 · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

When Bad Data Leads to Good Models OLMo: Accelerating the Science of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.164201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.164201Z digest=sha256:644a94e50e989bc4e4d7bbbd8961cf95373ed75962a5af1fcd97b0af745f5807

Observation ac0b6d1e-efde-43e6-9c69-4ecd9902cf7c · outbound

This paper cites Don't Stop Pretraining: Adapt Language Models to Domains and Tasks.

When Bad Data Leads to Good Models Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.169716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.169716Z digest=sha256:2bd544443633274e8f97ff81126d3cfe1ed6984443e5e4b9b9492b180907ea79

Observation 8547cafb-0b5f-4792-8aff-180dfcbc10d8 · outbound

This paper cites T oxi G en: A large-scale machine-generated dataset for adversarial and implicit hate speech detection.

When Bad Data Leads to Good Models T oxi G en: A large-scale machine-generated dataset for adversarial and implicit hate speech detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.175329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.175329Z digest=sha256:ffd9884231b8f710ab3ab7c70fbd4588cca39decec45f13fab380529df250d43

Observation a2e46bb3-c397-40f9-9707-44d402f05bdb · outbound

This paper cites Training Compute-Optimal Large Language Models.

When Bad Data Leads to Good Models Training Compute-Optimal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.180719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.180719Z digest=sha256:5b36333e7d250ea68593d0b93b1eb009cc193413ca75cb0542585da0abcf1da3

Observation 4b50b71e-24ee-41c8-8840-109b107aec9f · outbound

This paper cites Toxic comment classification challenge, 2018.

When Bad Data Leads to Good Models Toxic comment classification challenge, 2018

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.131842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.185733Z digest=sha256:e40d06b9f8055388d58f572bc017ab0ec4c89f94fea54230f6c29b0f95be4cfa

Observation bba9cde2-5b94-4b39-a4af-73f468a71afc · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

When Bad Data Leads to Good Models CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.191369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.191369Z digest=sha256:4cb3af302ab02350a472dc80b338a3e54f4d5cd8ddbe8b2083d304e225ab846a

Observation 70ac03cd-a255-48bc-80fe-72e4d5369f55 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

When Bad Data Leads to Good Models Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.196576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.196576Z digest=sha256:23029a9a5381fab2cdeca7f300579f9280763ba5b4cf4e299df4a3dfd4057069

Observation e469645b-796e-4199-9f17-d2c0fa4f5858 · outbound

This paper cites GeDi: Generative Discriminator Guided Sequence Generation.

When Bad Data Leads to Good Models GeDi: Generative Discriminator Guided Sequence Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.201730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.201730Z digest=sha256:b89b0da21d296d6c6e6b92ecece6753e9216629139abb7e4e2fc0ad4792f807a

Observation 2c6085e7-949a-426e-a80e-86e932d435e9 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

When Bad Data Leads to Good Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.206648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.206648Z digest=sha256:e770e89b32e9814c63861dade012f6c8b9c467b09c746346736c2f4149959eca

Observation 368755d9-9af7-4597-a1c6-31321658284f · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

When Bad Data Leads to Good Models Inference-time intervention: Eliciting truthful answers from a language model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.115476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.211713Z digest=sha256:46d244e192b410c45aca16793f5beb2db276b6d8c935e64efb0218f1a30b2ed4

Observation abf2596f-f907-4d30-bece-f075c645a25a · outbound

This paper cites Contrastive Decoding: Open-ended Text Generation as Optimization.

When Bad Data Leads to Good Models Contrastive Decoding: Open-ended Text Generation as Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.216849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.216849Z digest=sha256:3b4e787406462b582f3012059a9d369fa23629aa5489c0f8cdf43c0b77b9ea79

Observation 3b41d7ef-1f55-44ad-be48-c5cb2954f73f · outbound

This paper cites Disentangling transformer language models as superposed topic models.

When Bad Data Leads to Good Models Disentangling transformer language models as superposed topic models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.099144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.222381Z digest=sha256:8cffc58a9812704c20d30b6836baf07b26e58eff5ffc1481ac592bda8fb6c786

Observation 97e1c1e6-4f9b-4150-bf1d-582a8b7ebae1 · outbound

This paper cites Mitigating the Alignment Tax of RLHF.

When Bad Data Leads to Good Models Mitigating the Alignment Tax of RLHF

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.228033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.228033Z digest=sha256:b2230f9800610366fb08ba092c0ae342e90f66a710f5e7c6f83b8c3a7898e400

Observation 0c6791db-f08a-4e44-825b-e5ef10f89236 · outbound

This paper cites DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts.

When Bad Data Leads to Good Models DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.233785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.233785Z digest=sha256:af184ec59fd1482e0845159263cfa637a58a8f9b73b2da7dd97a4c69ae1bc78b

Observation c5157de0-8165-48ab-a585-33bab035de6d · outbound

This paper cites Amd-olmo: A series of 1b language models trained from scratch by amd on amd instinct™ mi250 gpus., October 2024.

When Bad Data Leads to Good Models Amd-olmo: A series of 1b language models trained from scratch by amd on amd instinct™ mi250 gpus., October 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.082439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.239167Z digest=sha256:264d7d9ebfe46fca53e49c1443129f243f26fa69c7e0e2f67b9154f08fdcfdcf

Observation c76c5b64-8033-4efa-ade9-0d5460a3e08b · outbound

This paper cites A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity.

When Bad Data Leads to Good Models A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.244162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.244162Z digest=sha256:ff7d02a5a2a7d81ef62f86b71e68f94d64f8abaf3a5ae776cebd8274c619d208

Observation 53856e84-5eda-4652-8888-4b64f1baa7d7 · outbound

This paper cites On linear representations and pretraining data frequency in language models.

When Bad Data Leads to Good Models On linear representations and pretraining data frequency in language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.066997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.249536Z digest=sha256:b89b32357f54d09151b51529964c318e18c0cddcf14be06bb2da4cb2a1812da9

Observation 8224ebc5-700f-4a65-9cc8-343bae33333d · outbound

This paper cites Linguistic regularities in continuous space word representations.

When Bad Data Leads to Good Models Linguistic regularities in continuous space word representations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.254124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.254124Z digest=sha256:988ad1ef68d01d53e7c1b69871912d237fdae9fd7228eca9ec3b01533ecf940f

Observation 44aa07b2-1ef2-42b2-af50-a93b04dabbb5 · outbound

This paper cites Interpreting gpt: The logit lens.

When Bad Data Leads to Good Models Interpreting gpt: The logit lens

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.039903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.259048Z digest=sha256:c6ea01afdbd563ab550a6ac8fbc036ebebd1577bf6cf14e66cbbd3790777a05f

Observation 613c008c-b05a-4f10-b2ad-5de148c0d92b · outbound

This paper cites Training language models to follow instructions with human feedback.

When Bad Data Leads to Good Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.264005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.264005Z digest=sha256:bb93410847ccea345825093cddddf0720a51ed1bc3deff1534b74c8127dce8f7

Observation f1188968-aaab-47db-9d52-c217652598ea · outbound

This paper cites Raiders of the lost kek: 3.5 years of augmented 4chan posts from the politically incorrect board.

When Bad Data Leads to Good Models Raiders of the lost kek: 3.5 years of augmented 4chan posts from the politically incorrect board

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:01.014366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.268529Z digest=sha256:9c3d7be3d5fb2466b4baf4c9e7e201757864bdf4b5ee0219c1b7fec508d0ed1f

Observation 913bea3b-fdc1-4e18-ab57-ed3e00371e55 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

When Bad Data Leads to Good Models Generative agents: Interactive simulacra of human behavior

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.273439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.273439Z digest=sha256:123d865c81cecca6f63c13256c85710163455e8a58213c0182e19a17bd1ab47b

Observation 854dec26-189e-492c-ab36-e59efe142155 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

When Bad Data Leads to Good Models The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.278006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.278006Z digest=sha256:851d44bb740d29c09d8af837c81aa965b2484d63a95d4464fe0e6c62b69718a4

Observation 175f1564-f108-4308-8292-d22005c637e2 · outbound

This paper cites Perspective | developers, 2024.

When Bad Data Leads to Good Models Perspective | developers, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:00.987592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.283369Z digest=sha256:1aa38de2fe27a616f02a27b88ca945fdf6ebb5173acf997ed0b1c61d578bcbcc

Observation bce706af-0f82-4858-ac05-332473f85003 · outbound

This paper cites Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models.

When Bad Data Leads to Good Models Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.288445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.288445Z digest=sha256:0be3efba10de398c0bf4498a8447c2051727cae68b78468a1cb9e1b8979c866a

Observation 91e7ccab-ec8a-4cab-8642-5dc0d9ae08ff · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

When Bad Data Leads to Good Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.293454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.293454Z digest=sha256:b1cb36d26612dffd3b3e12270626f2501ff12ce2e0088eaeb6d79232990d4657

Observation 916a22c0-021b-4236-bbf6-d9106ec40964 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

When Bad Data Leads to Good Models Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.298525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.298525Z digest=sha256:d45a77ddb367e07f22687965ff761d202ae0bf24307f6bd2c3b39a2d4e44b6c6

Observation cae6ff8d-8afd-4467-8cf0-e156b2b961c5 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

When Bad Data Leads to Good Models Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.303571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.303571Z digest=sha256:46e10eb903b3411d21d27c5e12346c620e68243a7a422f973a669005bfcbe568

Observation ee09727c-4ae3-466c-9222-be2923d71879 · outbound

This paper cites an unresolved cited work.

When Bad Data Leads to Good Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.308538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.308538Z digest=sha256:1f018b4b7f77fc355d7f8a949b6fffa432c1cc9c686bf4abf36a09b0f33cfe00

Observation ac047372-dd4c-4781-aff0-ebee555c821e · outbound

This paper cites Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp.

When Bad Data Leads to Good Models Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in nlp

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.313458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.313458Z digest=sha256:e7dc4e6c61d04945beb9ae092a005d3049efceb2af2a55ad2c1fdf1d6d734e22

Observation 7a8d70f4-356e-4be9-9c0c-1799478091f9 · outbound

This paper cites Process for adapting language models to society (palms) with values-targeted datasets.

When Bad Data Leads to Good Models Process for adapting language models to society (palms) with values-targeted datasets

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:00.952087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.318335Z digest=sha256:27cd557af81d32a12ff51bb82070763e0392432fa877fbd0128517516ec65619

Observation 3e9390c6-1218-43ac-8444-0c8c4253a16d · outbound

This paper cites BERT Rediscovers the Classical NLP Pipeline.

When Bad Data Leads to Good Models BERT Rediscovers the Classical NLP Pipeline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.322982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.322982Z digest=sha256:3953e858dc49e75f711df38c819eadea93f4c0ef09b9c3f0991542a21f4f0172

Observation 9dd612bc-acd7-4ed1-b92a-2794c86064c5 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

When Bad Data Leads to Good Models LaMDA: Language Models for Dialog Applications

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.327735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.327735Z digest=sha256:05278a00899450fb2f3557d35a7b865293d6f426247693ce19ec07319f19c4e2

Observation c966ef15-db07-43ee-bc29-deb4faecc115 · outbound

This paper cites Activation addition: Steering language models without optimization.

When Bad Data Leads to Good Models Activation addition: Steering language models without optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:00.936145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.332669Z digest=sha256:b42aeefe76a77a5e4a5aac6e289521668f57a6bbf5b3b1ab659a0c8754219d31

Observation e6672175-a16e-4f4e-a20f-4891429806aa · outbound

This paper cites Exploring the limits of domain-adaptive training for detoxifying large-scale language models.

When Bad Data Leads to Good Models Exploring the limits of domain-adaptive training for detoxifying large-scale language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:00.919423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.337413Z digest=sha256:2a19b77ad365b7b29cba46804044158bd36cfe772f1b457cec328216bfff046d

Observation 0ecbe4bc-f605-430b-9d7a-5a799b4d55b0 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

When Bad Data Leads to Good Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.342165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.342165Z digest=sha256:bec7546e551d401c035404431c6943f04fc2b5f9f69d0b82a60a29b9f8270c52

Observation 8e445eca-790a-49e6-87db-cbb605393df2 · outbound

This paper cites Lower bounds on the maximum cross correlation of signals (corresp.).

When Bad Data Leads to Good Models Lower bounds on the maximum cross correlation of signals (corresp.)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:25:00.902623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-15T23:25:00.347057Z digest=sha256:4640edfed2932cb5b393ef8db3d24d4d7ed3a496194ca1987f107d4e7ab0a7a9

Observation 6c75294e-4f85-42b7-9b59-344afc331f98 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

When Bad Data Leads to Good Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:25:00.352149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:25:00.352149Z digest=sha256:e83fadeba78b9090d8a803ab9edcab17bb17282cacdf4c43d6646c0d2e7ecd20

Pith citing papers

Observation 10b518e5-d95a-4060-9b57-3d13b862019b · inbound

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation cites this paper.

Making the Most of Limited Data: Score-Aware Training for Text-to-Music Generation When Bad Data Leads to Good Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:07:08.943080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T23:00:07.144841Z digest=sha256:9be2fd077e0e3b4ba3d5cc7e1366f6239cb0531b6ebaed1332d5eeabf661e21a

Observation e0b15acf-3aab-4fff-a4fe-63fb8b27140c · inbound

Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs cites this paper.

Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs When Bad Data Leads to Good Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:53:35.432279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:53:35.432279Z digest=sha256:fd7c714db964de13f05ff6a0dfe3d54c06ac25916802b689bae83d638a49a3be