Pith. sign in

Paper Citation Record · LEDGER

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2502.08657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08657 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:35.553049Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b90455ae-f6db-4334-8b9b-4dff47035811 · outbound

This paper cites GPT-4 Technical Report.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.351848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.351848Z digest=sha256:60e13676c496a6b45b139b2a568fd6f1759f4f0b7a0bca667e83242e7554a628

Observation b33bec3d-afa7-4659-a17d-3d8fe19b393a · outbound

This paper cites DeepSeek-V3 Technical Report.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions DeepSeek-V3 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.355831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.355831Z digest=sha256:04aff759fec30cc0c05227b18f24ed5f4298a9af19a30483bebd3aa5aa72df9a

Observation 44078d6b-9dfb-4df5-88da-a17a08d7b3e1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.359608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.359608Z digest=sha256:a303d162f030f3d0a0001007d1183cb7998b26c63bbadb2264c74811b0dc988d

Observation f07638cb-ec24-41fe-80bb-cdcd8e3a2928 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.363738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.363738Z digest=sha256:c452f1c452d2334173e0ca4a6a323c12ffbdb0a42ed8147b5ba610df7ecf14cf

Observation 34533a29-e8b8-41ee-9f1e-e620a028d73b · outbound

This paper cites The ai alignment problem: why it is hard, and where to start,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The ai alignment problem: why it is hard, and where to start,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.121774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.367673Z digest=sha256:82735a094bea0ee2b610512e4244cb991b911ff78f74877665e99a66c0ee6fec

Observation 0cc7659b-d836-47d5-9b15-275ffbaebbe6 · outbound

This paper cites Artificial intelligence, values, and alignment,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Artificial intelligence, values, and alignment,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.111997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.371153Z digest=sha256:ea2e4d05a5a8921ea649acd831dd6742df8ed0fef0f59dab7ffb37467c89b3e3

Observation 1c90bfd9-124d-4139-9ae0-dca01410a620 · outbound

This paper cites Self-instruct: Aligning language models with self- generated instructions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-instruct: Aligning language models with self- generated instructions,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.101679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.374888Z digest=sha256:270fa07887b346fd6fe31e58cd6f295fbda6e7c917b3dcaa3665ddc2368134cf

Observation 27b1f92f-53ec-4980-af67-70b609d027eb · outbound

This paper cites Self-alignment of large language models via monopolylogue-based social scene simulation,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-alignment of large language models via monopolylogue-based social scene simulation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.091151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.378320Z digest=sha256:0221b0439e1471ec1a737e7183778d4dede33e5db3be8661ee5d90173f2abd2b

Observation 648bc372-db72-460b-ba3d-ef72a90a9c22 · outbound

This paper cites Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.381498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.381498Z digest=sha256:06ee3234fb0212e65a51770f5bd8cc8d4ec21a9a9e9590056c1035653ef3df16

Observation d1901f4e-4d05-43f8-9993-53917d1a0576 · outbound

This paper cites Balancing differential privacy and utility: A relevance-based adaptive private fine-tuning framework for language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Balancing differential privacy and utility: A relevance-based adaptive private fine-tuning framework for language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.082574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.384930Z digest=sha256:47c914d1efb1b0b04a8e2b5f216d9177003adb689aeee42b092fe77fd260cdcd

Observation efcc71b5-2986-463c-80ce-ffc8f2c53bc7 · outbound

This paper cites Hard adversarial example mining for improving robust fairness,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Hard adversarial example mining for improving robust fairness,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.073695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.388365Z digest=sha256:3fdaa2dfc9f465c23aae253ebb7bc3b1eb2b088a88d16aaa8212bec80bc834e3

Observation 76e5cd17-b69e-40a1-b2d1-2e1afcbc0404 · outbound

This paper cites Finetuned language models are zero-shot learners,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Finetuned language models are zero-shot learners,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.064386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.392020Z digest=sha256:f130f324363986a1a42210714322785e2e3e0bc128d8f68ac3176ecd6704d688

Observation 805dd60b-66fe-4299-86d3-15c85574f5ed · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.395383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.395383Z digest=sha256:bbdfe7fff8824672decbb1d466036a717410d2c1130bc43db07d00a062312359

Observation 5bde28a1-a72a-4264-bcc1-2f435ef56ae6 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Fine-tuning aligned language models compromises safety, even when users do not intend to!

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.054791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.399198Z digest=sha256:cdc0247599a83c22476ae6398f74024c2bbe8a1175c9753dea1e00fd5ea357a3

Observation 39e23c0a-ca4b-435c-9abf-ce5dfc04f8f5 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Stanford alpaca: An instruction-following llama model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.403391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.403391Z digest=sha256:310f1d88d8c2dbe94d2bcdc2f12bda3e88276cf87ce844169ebfd3d719dfffd1

Observation 34156006-5528-4174-af5e-f2ab8bc7a871 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.406934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.406934Z digest=sha256:64934d9898fb2b45e310ab032649ecdc4297034c022effb3cb118efc3f303d88

Observation 00152d79-6299-4ce9-a8f6-718417e69234 · outbound

This paper cites Self-instruct: Aligning language models with self- generated instructions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-instruct: Aligning language models with self- generated instructions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.033196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.410273Z digest=sha256:8a77b43805a13b6e35c67bcfb2658b38c67cca8bfe9e4006712b996c37e18f92

Observation e2c4b06d-7662-4578-8b26-e836036081b2 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Training language models to follow instructions with human feedback,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.023714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.413493Z digest=sha256:85066024b0c126ce23850dcf552966012eda5f75a0775203420300b05e3a6272

Observation 8399b0d9-c134-414d-bce5-e7dddf93432e · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Beavertails: Towards improved safety alignment of llm via a human-preference dataset,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.014430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.417191Z digest=sha256:ddbbd9697695f4909d8e9d3025035a6d455c7e5f5b46166340790aedc1dd2ba5

Observation 1345c631-0cc0-4ecf-be07-6b52b9a6b7f5 · outbound

This paper cites Poisoning language models during instruction tuning,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Poisoning language models during instruction tuning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:36.005000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.420482Z digest=sha256:927eb91468f7fc32487286768b2f030e27701a0c70179b53ef5530ea95ac0d91

Observation fdbce315-085b-4fea-be85-68c1556a2737 · outbound

This paper cites Openassistant conversations-democratizing large language model align- ment,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Openassistant conversations-democratizing large language model align- ment,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.994753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.423742Z digest=sha256:7b5a82a9a435fcc5ccf4e111104afdecd81d3028c20cfeda312b1e2c69842df7

Observation 4b048045-507b-4655-9e95-31ed27700854 · outbound

This paper cites Principle-driven self-alignment of language models from scratch with minimal human supervision,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Principle-driven self-alignment of language models from scratch with minimal human supervision,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.985261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.427509Z digest=sha256:6c85e72889f4d5cdf073870aa21e37c4e45bbc6606ea660ec985731c0f0bf540

Observation d1dbda63-7d4d-4e17-a183-ab31d615a3df · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.430899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.430899Z digest=sha256:431c794b480a235199a48dddf500e14f78e10d95d1de049f30824ad9da201ce3

Observation 0da6d42f-2839-4d31-8723-13b76ff2c4d6 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.434529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.434529Z digest=sha256:01ad42361b7c5c2c28122e0ecd7c1bcdc89944d7f56e5a7e00ff1860eeefbd79

Observation 721bf31d-5100-4e67-830a-f891a999e833 · outbound

This paper cites Self-concept clarity development across the lifespan,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-concept clarity development across the lifespan,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.975277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.438293Z digest=sha256:cacfc19e743ef35f3c88128a9bd12b12103f187b6e40940c4111ae39202e7fe4

Observation 8cec7ea0-ce83-4d76-89ee-807450d21061 · outbound

This paper cites Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.441458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.441458Z digest=sha256:2f28d2131dd960a471d8b8ae9b001adcb47a74d6f3cb27c6d939a05bc9c918f9

Observation 67d604db-332f-4d2e-a985-bce45b67d65c · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.445137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.445137Z digest=sha256:59f425a119acfd2868e3093353c26703e4f2d4cd15044a4edb9b6c6686ff7316

Observation 8a8c2ee2-4adb-404e-af5b-88f0a39a02f2 · outbound

This paper cites Variational bayesian un- learning,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Variational bayesian un- learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.965834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.448132Z digest=sha256:0c64d91b8b40e9d34ae5b3d84820006d8f6b0161910bb6f8895758cb0438f282

Observation f09253eb-a71a-4520-8631-7bcd66c226f4 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.450781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.450781Z digest=sha256:8d8ff4a779a6d2ba46f5cf5d11eec02546b1f6b363c8855b4756d53de36dae5c

Observation 4f879638-9652-4351-baca-4f6a1d159684 · outbound

This paper cites A Survey of Large Language Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions A Survey of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.453762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.453762Z digest=sha256:d517171bff981d3a11764771e5c527764742cc42c4ac740b4f4725ec566940c1

Observation 92cdd58e-ad1c-4c04-8397-27b7704acb45 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions On the Opportunities and Risks of Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.456785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.456785Z digest=sha256:c987fe110f7199eed4b6d197a3e48e60a75cea50b0b5c8b24a1c89390f15bcf7

Observation 55eb9564-193f-4d61-aa02-4d7628961252 · outbound

This paper cites Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.955845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.460211Z digest=sha256:a81449c3a9633c05a00a852981d7e048f2c614f2cc268a968b172ef2a624ddc6

Observation f9b95246-da35-47f7-a07a-4b72592d8596 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Fine-grained human feedback gives better rewards for language model training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.944549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.463441Z digest=sha256:741b8b1671196e69aee9c3e085aa1403c18e7b3fa5076a66e76d6d147b7f06f1

Observation c5f597f0-13b9-422e-87bb-785a3a9c9b4e · outbound

This paper cites Learning to summarize with human feedback,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Learning to summarize with human feedback,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.933971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.466743Z digest=sha256:c88835731e8f5427c8562e8b0af00ccdfa77dd11db955a208eea7981be7498bb

Observation 57659805-3e28-4c17-a855-6d5f70f22100 · outbound

This paper cites Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.925355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.470155Z digest=sha256:f928ec870e665341af2e471b1fbfc0240c9bd111ad40bc9965181db3fb7d2515

Observation 77cd4e56-c207-465f-b4b4-0d46f6b51d63 · outbound

This paper cites Gaining wisdom from setbacks: Aligning large language models via mistake analysis,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Gaining wisdom from setbacks: Aligning large language models via mistake analysis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.916201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.473313Z digest=sha256:b566f809d5c93618fa6183668a6e05fe62a83b24de5647af6ea3fcdfa205bc92

Observation 91041920-d5a2-45a2-8069-8d8b760dae0f · outbound

This paper cites Understanding negative samples in instance dis- criminative self-supervised representation learning,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Understanding negative samples in instance dis- criminative self-supervised representation learning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.907055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.476777Z digest=sha256:3ed622f7423892862119f2e3f48dff5eb05d6e91527f3ff2db5b167940862524

Observation c4f2e0d2-7b2f-4c36-a474-aa7acc360373 · outbound

This paper cites Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.479808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.479808Z digest=sha256:8ae566c0a5e2bd69b92b6ce66f616edb5c5d2dd71e078b23e06f796c4445fb18

Observation 00f9962e-41d9-4b82-8d17-eebc325c02ee · outbound

This paper cites What makes for good views for contrastive learning?.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions What makes for good views for contrastive learning?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.483473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.483473Z digest=sha256:c6997e542f7ba4a6a82b6863c1e2d383ed68af5597aa31064d4d8330195bc553

Observation 6ea4fda9-0edd-4acf-b225-23184ce5552e · outbound

This paper cites Neural text degeneration with unlikelihood training,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Neural text degeneration with unlikelihood training,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.891016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.486507Z digest=sha256:9c8fbd01a47479ed566317dcbb6307ed96322d2c5b5969068daa7d05df292bc8

Observation 4cde7402-0787-42a0-8e7b-cf85aa83ef18 · outbound

This paper cites Openai. gpt-4v(ision) system card,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Openai. gpt-4v(ision) system card,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.881487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.489654Z digest=sha256:5fc29b0a3d07ce2494bf2d94f59d4a1d942f866b7edc80f2a9e7613283fa6975

Observation b78a06e6-75c4-4d4a-9fb6-f9ec1518778e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.492858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.492858Z digest=sha256:2bdc92706210b2de725865fe76a267750a37b87da9e4d9caa7c11c2a4c2b0193

Observation 2c835252-f39b-495d-b08e-1a973428ef69 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.871580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.496551Z digest=sha256:7c0580ee55ef013c824de6ce381e325af87434db5c0bb32ae95a0e0d54731e7b

Observation a4bbf00f-7f71-48e1-8023-4e5969b47e7e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Gemini: A Family of Highly Capable Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.499871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.499871Z digest=sha256:608a779ababa35a8536718364886a8699a3456ca965e55afa3a657909d1f8761

Observation b9118a71-173d-4d90-94af-3407ca079fd2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Chain-of-thought prompting elicits reasoning in large language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.503523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.503523Z digest=sha256:d8e8beecd5c6e0bd2c8ba5551da192777c2185effe7f4b3eb5bf3ef63c932aca

Observation 23e369c3-388c-4625-afc6-ab1058cce931 · outbound

This paper cites The wisdom of hindsight makes language models better instruction followers,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The wisdom of hindsight makes language models better instruction followers,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.855924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.506641Z digest=sha256:a9a1da131a8f2d65f50a4d9c54babd0abe1ed9c0891e01cb11524692b1820dec

Observation 340786a3-7012-4239-9582-dabbf100e490 · outbound

This paper cites Koala: A dialogue model for academic research,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Koala: A dialogue model for academic research,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.846317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.509806Z digest=sha256:d40b4f90d18b88e57a4fb938b0d0eb6d9e8a633e99c15ede8f585cc5352b2c31

Observation 61f8cb44-e1ea-464a-900f-ba2eac9be164 · outbound

This paper cites Glm-130b: An open bilingual pre-trained model,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Glm-130b: An open bilingual pre-trained model,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.836713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.513267Z digest=sha256:4cbc9b3320f6050ae492a54dcc3433431488dde7f1286d27d8de993152aaae8b

Observation a4e281db-3fdd-4ff6-ada7-49ab27538497 · outbound

This paper cites The curious case of neural text degeneration,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The curious case of neural text degeneration,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.826185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.516447Z digest=sha256:a47c4873b1b4d208090679baed0180748937787ee4e0b1c513fc9c75618ce89a

Observation 0e9c7d08-786e-4f8e-9bac-9a6946ca5f81 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Lora: Low-rank adaptation of large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.815745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.519585Z digest=sha256:11326c9aef09cefea9d4971dd832153c054a18b3b1c2fa91fc39822d785c6dda

Observation 9ef50156-efc3-498c-a623-37771202d5ad · outbound

This paper cites The Llama 3 Herd of Models.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The Llama 3 Herd of Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.523090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.523090Z digest=sha256:020eda9cb8187ba3f1c1c9074fc22532a430275c08e2145b40f14a12c7518dd4

Observation 3fdce5e9-8d92-4dc6-9d1d-0bf8844b6519 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.805988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.526699Z digest=sha256:077e56a0844a32575373e7ec9000945132b778070f16c221df25ee81049c093b

Observation dd711585-734e-43ee-8ba3-4e08b50fd056 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Autodan: Generating stealthy jailbreak prompts on aligned large language models,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.796019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.529960Z digest=sha256:3ab10e3c678622d2282efd626c74156262f5a6d9cabd7e7a4cf481dd0bc13254

Observation cf6c2c09-2ba9-4856-9d5f-4b273dcf9166 · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.786305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.533362Z digest=sha256:6ed6376161e3342e0c52d52349ad1dc07abc37d2ac54fc4b62c26dd3a3c02872

Observation cf95be37-1915-4c0a-8207-3ac6d74ec4e3 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Truthfulqa: Measuring how models mimic human falsehoods,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.776407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.536655Z digest=sha256:cf9981dfcebc8d80fdbcca37a626ecf26bbd9ed94dd8cc1600b8c1a53fd3390d

Observation 937b1e5a-7f3e-4358-b548-0f81431ab153 · outbound

This paper cites Measuring massive multitask language understanding,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Measuring massive multitask language understanding,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.539940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.539940Z digest=sha256:576b9d5f2febd23f35f66039ec3e3a8c74b72ea1112c58ebd8d8d39a90278999

Observation b23b4cd2-24f3-4aec-a5ae-6609841d685b · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Social iqa: Commonsense reasoning about social interactions,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.761578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.543293Z digest=sha256:64293f52b10ce02e0fd3cf2a32df37f4acd07c706a14f0b36ec338f5b61822de

Observation 5534bcea-024b-408b-bde5-a85d0ea447b1 · outbound

This paper cites Mitigating the alignment tax of rlhf,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Mitigating the alignment tax of rlhf,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:35.750924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:07:35.546791Z digest=sha256:98843b1c618b03cf358c4aca83f044fd2cc49ee887d19c882fc8f1a4a880dfcd

Observation 712610c3-5c13-4414-9edf-2fd41067a154 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Direct preference optimization: Your language model is secretly a reward model,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.550161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.550161Z digest=sha256:665a84b1ef6e4143e58e9d3c736cc2958e55a960846e45d084e4b005ad387068

Observation 1bfb14e8-30c3-4297-bb0d-779945f7f404 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions KTO: Model Alignment as Prospect Theoretic Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.553049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.553049Z digest=sha256:62c638c34363e46e1f89c21fb8619464a03b62590672f3aca1f83cb8f8fd8c38

Pith citing papers

No inbound Pith citation observations are available.