Pith. sign in

Paper Citation Record · LEDGER

Rule Based Rewards for Language Model Safety

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2411.01111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.01111 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:27:33.269702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:10:42.078748Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3d8be13-5f55-4c14-9033-84d6af8a8726 · inbound

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning cites this paper.

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:27:33.269702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:27:33.269702Z digest=sha256:5fd87ffdbe92af10c07c14a67625793026efa6427f5d388668d7c59eedcac5dc

Observation 7be1af07-4abb-4b01-a920-0ab8dbe4f85d · inbound

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI cites this paper.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Rule Based Rewards for Language Model Safety

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.325631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.325631Z digest=sha256:8503e876e243903ad6faaa6bd179e440d2963f6d6a11930c8bf4ab01ea77e62f

Observation 4b4a198e-3d01-452e-a5cb-12c5f2dcbb0a · inbound

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs cites this paper.

Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs Rule Based Rewards for Language Model Safety

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.367853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:58:47.367853Z digest=sha256:fd69eaaf327bc6311e7769b6139c543afef1b08d88fdbc753c4e8e2419760daf

Observation 5712b820-e111-4872-9133-134ee783c7dd · inbound

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models cites this paper.

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models Rule Based Rewards for Language Model Safety

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T18:05:52.323811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:05:52.323811Z digest=sha256:e09f2d9efc1011d6afdb51edde8c25367a7e5834b9e716ffd5aeedaf2acfc712

Observation 0f2d1d33-7747-4a13-b729-87eacef238f0 · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines Rule Based Rewards for Language Model Safety

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:35.865028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:35.865028Z digest=sha256:f332abe7a83d556be4cd2e5caa090582a773c89f8e2603aad4093ba69280f7f3

Observation 1be1e9ba-9aab-4af3-9a66-735e7c64dbff · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rule Based Rewards for Language Model Safety

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.390130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.390130Z digest=sha256:2b3b95d68695d67e14c8b5a7aa7872f028fdd9824bd97c04308ad84b0f265c80

Observation 17c71015-1987-4a27-a882-f6481b4d69cc · inbound

Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition cites this paper.

Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition Rule Based Rewards for Language Model Safety

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:48.267448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:40:48.267448Z digest=sha256:a241f369a894f2a38da19f336c56d7c138e6169fca312b1aa526fceed591761e

Observation ad1c8170-f1f3-4b98-b887-a95067ce60f6 · inbound

Saffron-1: Safety Inference Scaling cites this paper.

Saffron-1: Safety Inference Scaling Rule Based Rewards for Language Model Safety

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.867442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.867442Z digest=sha256:146c35c1a9058238704f3578f7ce533c10137ee3f02427b9c546469d3a4de371

Observation 2295d038-75a9-4a7b-92da-5174726ff034 · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.427164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.427164Z digest=sha256:92bdb27b5b68769a9f350b07e720fe878e35c6d174d28f65d748c5afcc537592

Observation b33cbaa3-597c-48cf-a2b5-611d108b06ad · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Rule Based Rewards for Language Model Safety

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:38.316889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:38.316889Z digest=sha256:bcff9c897cd13061bbb48bc75de26d35f6939d1265f337a2c28c762151e848ff

Observation 5a377456-02a8-4df1-997f-7ee0d613ce03 · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Rule Based Rewards for Language Model Safety

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:56.071946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:56.071946Z digest=sha256:8db5bfc25b08e53a7983f6bc0d26a9d41e44607a2c4bd00225e16dedd3024550

Observation bb7de678-5cf5-4827-82b0-70da4b9c2303 · inbound

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards cites this paper.

RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards Rule Based Rewards for Language Model Safety

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.080709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T22:09:47.649346Z digest=sha256:837a52b7a16ad7f235b1c1a1b9d4353a7ca92d0579f718ac64eb328ee28d35c8

Observation 9af86ce4-e33d-4c79-8b70-148c9cabb665 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Rule Based Rewards for Language Model Safety

Reference 251

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:e6a576e3b9bc3c54616d9040149785a959f9209ee7934024f18a85b9257de71a

Observation 7deb59a2-68fe-428c-a488-23b16ef433e5 · inbound

Mach-Mind-4-Flash Technical Report cites this paper.

Mach-Mind-4-Flash Technical Report Rule Based Rewards for Language Model Safety

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T03:29:34.486347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:29:34.486347Z digest=sha256:877f8e772ca4f0112f334f1026d1adfae1e338abd6ab51fd42db2e1f3db57309

Observation 219d03dd-1472-40de-bdd5-1ad7d8a314ea · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Rule Based Rewards for Language Model Safety

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.055928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:26.055928Z digest=sha256:98d8107a4a2ddcb3eda6197223222edb6202594ef0510656424f328f7be2c9cd