Pith. sign in

Paper Citation Record · LEDGER

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2502.08657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08657 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:35.553049Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b90455ae-f6db-4334-8b9b-4dff47035811 · outbound

This paper cites GPT-4 Technical Report.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.351848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.351848Z digest=sha256:60e13676c496a6b45b139b2a568fd6f1759f4f0b7a0bca667e83242e7554a628

Observation b33bec3d-afa7-4659-a17d-3d8fe19b393a · outbound

This paper cites DeepSeek-V3 Technical Report.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions DeepSeek-V3 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.355831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.355831Z digest=sha256:04aff759fec30cc0c05227b18f24ed5f4298a9af19a30483bebd3aa5aa72df9a

Observation 44078d6b-9dfb-4df5-88da-a17a08d7b3e1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.359608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.359608Z digest=sha256:a303d162f030f3d0a0001007d1183cb7998b26c63bbadb2264c74811b0dc988d

Observation f07638cb-ec24-41fe-80bb-cdcd8e3a2928 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.363738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.363738Z digest=sha256:c452f1c452d2334173e0ca4a6a323c12ffbdb0a42ed8147b5ba610df7ecf14cf

Observation 34533a29-e8b8-41ee-9f1e-e620a028d73b · outbound

This paper cites The ai alignment problem: why it is hard, and where to start,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The ai alignment problem: why it is hard, and where to start,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.121774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.367673Z digest=sha256:135264daa69eb152dd57b84980da7e3dc4f292e4ef83b2636558119dcdfb360c

Observation 0cc7659b-d836-47d5-9b15-275ffbaebbe6 · outbound

This paper cites Artificial intelligence, values, and alignment,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Artificial intelligence, values, and alignment,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.111997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.371153Z digest=sha256:524d0a9fdd3b4728deaf12c0d1cb38d2b21377ba7683b7a6b9266e12266b49f0

Observation 1c90bfd9-124d-4139-9ae0-dca01410a620 · outbound

This paper cites Self-instruct: Aligning language models with self- generated instructions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-instruct: Aligning language models with self- generated instructions,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.101679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.374888Z digest=sha256:4772ee13c2829e6715b08743905c012978775ffb246c1b51e0841bb0a044c9ea

Observation 27b1f92f-53ec-4980-af67-70b609d027eb · outbound

This paper cites Self-alignment of large language models via monopolylogue-based social scene simulation,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-alignment of large language models via monopolylogue-based social scene simulation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.091151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.378320Z digest=sha256:178cfb40cabfda68228939f846330ad9ff714f1d21c1a299f596638323d23d8a

Observation 648bc372-db72-460b-ba3d-ef72a90a9c22 · outbound

This paper cites Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.381498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.381498Z digest=sha256:06ee3234fb0212e65a51770f5bd8cc8d4ec21a9a9e9590056c1035653ef3df16

Observation d1901f4e-4d05-43f8-9993-53917d1a0576 · outbound

This paper cites Balancing differential privacy and utility: A relevance-based adaptive private fine-tuning framework for language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Balancing differential privacy and utility: A relevance-based adaptive private fine-tuning framework for language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.082574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.384930Z digest=sha256:f64d50365d94cb7846df64610fa054d59818fcf48147a40d300b586dd9ae715e

Observation efcc71b5-2986-463c-80ce-ffc8f2c53bc7 · outbound

This paper cites Hard adversarial example mining for improving robust fairness,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Hard adversarial example mining for improving robust fairness,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.073695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.388365Z digest=sha256:4220490c2a07cf0707011112fb6cd1efd9b815c199b0ac998606dd46e267d6c5

Observation 76e5cd17-b69e-40a1-b2d1-2e1afcbc0404 · outbound

This paper cites Finetuned language models are zero-shot learners,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Finetuned language models are zero-shot learners,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.064386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.392020Z digest=sha256:c874e049ecb5f4845b2e8ae357f3fe1c88bd5023f55a81d66b4393ab03ef4a20

Observation 805dd60b-66fe-4299-86d3-15c85574f5ed · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.395383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.395383Z digest=sha256:bbdfe7fff8824672decbb1d466036a717410d2c1130bc43db07d00a062312359

Observation 5bde28a1-a72a-4264-bcc1-2f435ef56ae6 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Fine-tuning aligned language models compromises safety, even when users do not intend to!

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.054791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.399198Z digest=sha256:d74ed3bf39928c39e9a0beeedbd14557c89c9918580338611f4b125820c46205

Observation 39e23c0a-ca4b-435c-9abf-ce5dfc04f8f5 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Stanford alpaca: An instruction-following llama model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.403391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.403391Z digest=sha256:310f1d88d8c2dbe94d2bcdc2f12bda3e88276cf87ce844169ebfd3d719dfffd1

Observation 34156006-5528-4174-af5e-f2ab8bc7a871 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.406934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.406934Z digest=sha256:64934d9898fb2b45e310ab032649ecdc4297034c022effb3cb118efc3f303d88

Observation 00152d79-6299-4ce9-a8f6-718417e69234 · outbound

This paper cites Self-instruct: Aligning language models with self- generated instructions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-instruct: Aligning language models with self- generated instructions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.033196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.410273Z digest=sha256:c3ce29eb2e97532954cb5a76187acaf337e637319fdedc2fd35e35074032a445

Observation e2c4b06d-7662-4578-8b26-e836036081b2 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Training language models to follow instructions with human feedback,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.023714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.413493Z digest=sha256:84d151d5a4eb8847ae04130b0ae749ef6285654093238e33f606597f918ea41a

Observation 8399b0d9-c134-414d-bce5-e7dddf93432e · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Beavertails: Towards improved safety alignment of llm via a human-preference dataset,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.014430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.417191Z digest=sha256:cda8daced6224398185c2bc76e9145b993d1b0947ab4ddccd527d1e16de504ce

Observation 1345c631-0cc0-4ecf-be07-6b52b9a6b7f5 · outbound

This paper cites Poisoning language models during instruction tuning,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Poisoning language models during instruction tuning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.005000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.420482Z digest=sha256:78f0b3435d31e85ecce769ed09715f692e666ad5afe19144fc3a0a60112d6f96

Observation fdbce315-085b-4fea-be85-68c1556a2737 · outbound

This paper cites Openassistant conversations-democratizing large language model align- ment,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Openassistant conversations-democratizing large language model align- ment,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.994753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.423742Z digest=sha256:39f81d454d98fb620db0219aa541acbeebed759ebf0abb636120f03db49e7ea8

Observation 4b048045-507b-4655-9e95-31ed27700854 · outbound

This paper cites Principle-driven self-alignment of language models from scratch with minimal human supervision,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Principle-driven self-alignment of language models from scratch with minimal human supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.985261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.427509Z digest=sha256:acaeb5d1ae359ede7b047d0c267e967bed4d64c3a767e499d319352ad265f481

Observation d1dbda63-7d4d-4e17-a183-ab31d615a3df · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.430899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.430899Z digest=sha256:431c794b480a235199a48dddf500e14f78e10d95d1de049f30824ad9da201ce3

Observation 0da6d42f-2839-4d31-8723-13b76ff2c4d6 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.434529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.434529Z digest=sha256:01ad42361b7c5c2c28122e0ecd7c1bcdc89944d7f56e5a7e00ff1860eeefbd79

Observation 721bf31d-5100-4e67-830a-f891a999e833 · outbound

This paper cites Self-concept clarity development across the lifespan,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-concept clarity development across the lifespan,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.975277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.438293Z digest=sha256:2dc9d9bfb222cd1d55ee383c380a47aa130af222ea1d6fbaf6deccc3fc8c83dc

Observation 8cec7ea0-ce83-4d76-89ee-807450d21061 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.441458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.441458Z digest=sha256:2f28d2131dd960a471d8b8ae9b001adcb47a74d6f3cb27c6d939a05bc9c918f9

Observation 67d604db-332f-4d2e-a985-bce45b67d65c · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.445137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.445137Z digest=sha256:59f425a119acfd2868e3093353c26703e4f2d4cd15044a4edb9b6c6686ff7316

Observation 8a8c2ee2-4adb-404e-af5b-88f0a39a02f2 · outbound

This paper cites Variational bayesian un- learning,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Variational bayesian un- learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.965834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.448132Z digest=sha256:efa6f2af7cde2d3bf8d5b3dddcbb062458959afdd76ce9772eefb2371c98bbdc

Observation f09253eb-a71a-4520-8631-7bcd66c226f4 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.450781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.450781Z digest=sha256:8d8ff4a779a6d2ba46f5cf5d11eec02546b1f6b363c8855b4756d53de36dae5c

Observation 4f879638-9652-4351-baca-4f6a1d159684 · outbound

This paper cites A Survey of Large Language Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions A Survey of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.453762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.453762Z digest=sha256:d517171bff981d3a11764771e5c527764742cc42c4ac740b4f4725ec566940c1

Observation 92cdd58e-ad1c-4c04-8397-27b7704acb45 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions On the Opportunities and Risks of Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.456785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.456785Z digest=sha256:c987fe110f7199eed4b6d197a3e48e60a75cea50b0b5c8b24a1c89390f15bcf7

Observation 55eb9564-193f-4d61-aa02-4d7628961252 · outbound

This paper cites Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.955845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.460211Z digest=sha256:cfc545234f0c3c3c51fd00ac053eb1937cb1388ed5737e77213009e5d353e034

Observation f9b95246-da35-47f7-a07a-4b72592d8596 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Fine-grained human feedback gives better rewards for language model training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.944549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.463441Z digest=sha256:619cfac39d2b250acacea1ef731ba2e91c351c4f92b4da11252188f55fe51725

Observation c5f597f0-13b9-422e-87bb-785a3a9c9b4e · outbound

This paper cites Learning to summarize with human feedback,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Learning to summarize with human feedback,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.933971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.466743Z digest=sha256:c06ef2a9a9e7f628a45341f86ba750dffe8784ee8ec9753ad2187eae96a5f75b

Observation 57659805-3e28-4c17-a855-6d5f70f22100 · outbound

This paper cites Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.925355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.470155Z digest=sha256:b9a447a19c1dbb13c77cf0b6137eea9f73eb2106e7445df2a22f5916b6435742

Observation 77cd4e56-c207-465f-b4b4-0d46f6b51d63 · outbound

This paper cites Gaining wisdom from setbacks: Aligning large language models via mistake analysis,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Gaining wisdom from setbacks: Aligning large language models via mistake analysis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.916201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.473313Z digest=sha256:471d231a6cef5589cbdb7d9f500932dda4676cb630dbef1f815cb9acc65fcf49

Observation 91041920-d5a2-45a2-8069-8d8b760dae0f · outbound

This paper cites Understanding negative samples in instance dis- criminative self-supervised representation learning,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Understanding negative samples in instance dis- criminative self-supervised representation learning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.907055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.476777Z digest=sha256:5f5f65973230333a6a58537251f9d2f1fc8ec2083852e010b299954d1bfbd745

Observation c4f2e0d2-7b2f-4c36-a474-aa7acc360373 · outbound

This paper cites Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.479808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.479808Z digest=sha256:8ae566c0a5e2bd69b92b6ce66f616edb5c5d2dd71e078b23e06f796c4445fb18

Observation 00f9962e-41d9-4b82-8d17-eebc325c02ee · outbound

This paper cites What makes for good views for contrastive learning?.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions What makes for good views for contrastive learning?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.483473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.483473Z digest=sha256:c6997e542f7ba4a6a82b6863c1e2d383ed68af5597aa31064d4d8330195bc553

Observation 6ea4fda9-0edd-4acf-b225-23184ce5552e · outbound

This paper cites Neural text degeneration with unlikelihood training,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Neural text degeneration with unlikelihood training,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.891016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.486507Z digest=sha256:b75c638515dfdffc7098add21500cf3ac1f47a4036f828545b13752e5cec7c6c

Observation 4cde7402-0787-42a0-8e7b-cf85aa83ef18 · outbound

This paper cites Openai. gpt-4v(ision) system card,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Openai. gpt-4v(ision) system card,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.881487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.489654Z digest=sha256:98d70b6a9b81953c90b4c93a809d9f65758441c3d6fb6cac7b6cd757375e3b0f

Observation b78a06e6-75c4-4d4a-9fb6-f9ec1518778e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.492858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.492858Z digest=sha256:2bdc92706210b2de725865fe76a267750a37b87da9e4d9caa7c11c2a4c2b0193

Observation 2c835252-f39b-495d-b08e-1a973428ef69 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.871580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.496551Z digest=sha256:e06f3f92425cbedae3b9a368b3e6465216e6bbaccd44db2e9c5680f2ef784fc7

Observation a4bbf00f-7f71-48e1-8023-4e5969b47e7e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Gemini: A Family of Highly Capable Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.499871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.499871Z digest=sha256:608a779ababa35a8536718364886a8699a3456ca965e55afa3a657909d1f8761

Observation b9118a71-173d-4d90-94af-3407ca079fd2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Chain-of-thought prompting elicits reasoning in large language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.503523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.503523Z digest=sha256:d8e8beecd5c6e0bd2c8ba5551da192777c2185effe7f4b3eb5bf3ef63c932aca

Observation 23e369c3-388c-4625-afc6-ab1058cce931 · outbound

This paper cites The wisdom of hindsight makes language models better instruction followers,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The wisdom of hindsight makes language models better instruction followers,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.855924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.506641Z digest=sha256:9933327ef739fb7fc36ed15fe39670fbd6ca0237c439586a43eb5cc42683c9bf

Observation 340786a3-7012-4239-9582-dabbf100e490 · outbound

This paper cites Koala: A dialogue model for academic research,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Koala: A dialogue model for academic research,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.846317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.509806Z digest=sha256:b81b514b41c4fc91a3fe7b669a8b825bbbac60d244b3c94b7e45ed38924abb2f

Observation 61f8cb44-e1ea-464a-900f-ba2eac9be164 · outbound

This paper cites Glm-130b: An open bilingual pre-trained model,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Glm-130b: An open bilingual pre-trained model,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.836713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.513267Z digest=sha256:2037a04258471b9a22afe725d471e77f85c70692a21765bb113b746d21884dd5

Observation a4e281db-3fdd-4ff6-ada7-49ab27538497 · outbound

This paper cites The curious case of neural text degeneration,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The curious case of neural text degeneration,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.826185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.516447Z digest=sha256:d02bf1fd6a7d1cf5b74b21e42db564fdc622eca0ae4a31f2891df76d464acd7f

Observation 0e9c7d08-786e-4f8e-9bac-9a6946ca5f81 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Lora: Low-rank adaptation of large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.815745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.519585Z digest=sha256:395146796d9be18c2f934efe46720453cd61824b4cf8e124ff006c3496738f58

Observation 9ef50156-efc3-498c-a623-37771202d5ad · outbound

This paper cites The Llama 3 Herd of Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The Llama 3 Herd of Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.523090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.523090Z digest=sha256:020eda9cb8187ba3f1c1c9074fc22532a430275c08e2145b40f14a12c7518dd4

Observation 3fdce5e9-8d92-4dc6-9d1d-0bf8844b6519 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.805988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.526699Z digest=sha256:ebff1a83d562308e5b9ca16fbf3ca758517b5c2244da186612c7315c60a4dea3

Observation dd711585-734e-43ee-8ba3-4e08b50fd056 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Autodan: Generating stealthy jailbreak prompts on aligned large language models,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.796019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.529960Z digest=sha256:fd0b0397a4391cf44a91197a05f4283c450935fd3d926d70cc49800def5e9d8d

Observation cf6c2c09-2ba9-4856-9d5f-4b273dcf9166 · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.786305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.533362Z digest=sha256:14d478a4b852eafbed10796dcea306b376fb6c8bb9025c1d486b173ba13ee108

Observation cf95be37-1915-4c0a-8207-3ac6d74ec4e3 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Truthfulqa: Measuring how models mimic human falsehoods,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.776407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.536655Z digest=sha256:4565ec38f6e4dcbcf6f776494f0ac23d79a8749d891123368272460da359b97f

Observation 937b1e5a-7f3e-4358-b548-0f81431ab153 · outbound

This paper cites Measuring massive multitask language understanding,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Measuring massive multitask language understanding,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.539940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.539940Z digest=sha256:576b9d5f2febd23f35f66039ec3e3a8c74b72ea1112c58ebd8d8d39a90278999

Observation b23b4cd2-24f3-4aec-a5ae-6609841d685b · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Social iqa: Commonsense reasoning about social interactions,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.761578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.543293Z digest=sha256:2245de857c29710647dda75e19ff674c574a0c1d9e99a1f9560a19a1194278dc

Observation 5534bcea-024b-408b-bde5-a85d0ea447b1 · outbound

This paper cites Mitigating the alignment tax of rlhf,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Mitigating the alignment tax of rlhf,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.750924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:07:35.546791Z digest=sha256:236dc0b2f2d8b836c3304cc4306fd6017fd2df20822ba5fc048eb2a794067899

Observation 712610c3-5c13-4414-9edf-2fd41067a154 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Direct preference optimization: Your language model is secretly a reward model,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.550161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.550161Z digest=sha256:665a84b1ef6e4143e58e9d3c736cc2958e55a960846e45d084e4b005ad387068

Observation 1bfb14e8-30c3-4297-bb0d-779945f7f404 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions KTO: Model Alignment as Prospect Theoretic Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.553049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.553049Z digest=sha256:62c638c34363e46e1f89c21fb8619464a03b62590672f3aca1f83cb8f8fd8c38

Pith citing papers

No inbound Pith citation observations are available.